For years, one of the most common questions about medical AI has been:
How accurate are its answers?
But as AI moves closer to real clinical settings, another question is becoming just as important:
When the system does not know, can it choose not to answer?
A recent study published in npj Digital Medicine evaluated EyeSeek, a large language model system designed for primary eye care.
One particularly important feature of the system is its ability to recognize situations in which its internal knowledge may be insufficient. Instead of continuing to generate a confident-sounding answer, the system can abstain and seek external information.
This may sound like a small technical feature.
In healthcare, it could represent an important shift.
The Problem Is Not That AI Sometimes Does Not Know
The more serious problem is:
AI may not know — but still sound as if it does.
In an ordinary conversation, an incorrect answer may simply create confusion.
In healthcare, however, incorrect information may influence whether someone seeks medical attention, delays care, undergoes further testing, or misunderstands their health risk.
A mature medical AI system may therefore need more than the ability to answer questions.
It may also need the ability to determine:
Do I have enough evidence to answer this question?
This moves medical AI beyond answer generation toward a more fundamental capability:
uncertainty management.
“I Don’t Know” Can Be a Capability
The EyeSeek research also included a small, single-center prospective real-world study.
There were 84 participants in the AI-assisted group and 86 in the non-assisted group. AI assistance was associated with higher referral adherence and improvement in health literacy.
These findings are encouraging, but the evidence boundary is important.
The study was relatively small and conducted at a single center. It does not establish that this approach will improve ophthalmic care across different healthcare systems, populations, or clinical environments.
The more important idea may be the system design itself:
Medical AI does not necessarily need to answer every question.
In a high-stakes environment, an AI system that can recognize its limits, abstain when appropriate, and hand a question back to a clinician may ultimately be more useful than one designed to answer everything.
From Triage to Knowing When to Stop
This question becomes even more interesting when viewed alongside recent research on AI-assisted medical triage.
Triage asks AI to move beyond:
“What might this condition be?”
toward:
“What should this person do next?”
Should the patient remain at home?
Arrange a medical appointment?
Seek urgent assessment?
Once AI begins influencing actions rather than simply providing information, uncertainty becomes much more consequential.
The next generation of medical AI may therefore need to answer four different questions:
When can I answer?
When can I recommend an action?
When should I acknowledge uncertainty?
When should I hand the decision back to a healthcare professional?
These questions may ultimately matter as much as improvements in model accuracy.
Regulatory Approval Is Not the End of the Evidence Question
Another recent npj Digital Medicine study examined 77 AI diagnostic software products in pathology and hematologic morphology that had received CE marking or FDA authorization.
Researchers identified publicly available, device-specific performance evidence for 30 of the 77 products. For 47 products, they were unable to identify such publicly accessible performance evidence.
An important distinction is necessary:
This does not mean those products had “no evidence.”
Regulators may have reviewed evidence that was not readily available in the public literature.
Instead, the study raises a different question:
After an AI medical device receives regulatory authorization, how much of its performance evidence can clinicians, researchers, and the public actually examine?
The study also identified substantial variation in the performance metrics reported across products, making direct comparisons more difficult.
As medical AI becomes more common, evaluation may therefore need to move beyond:
“Has it been authorized?”
to additional questions:
How transparent is the evidence?
Can different systems be meaningfully compared?
How does the technology perform in real-world clinical practice?
Diagnostics Must Ultimately Lead to Action
A related development is occurring in Alzheimer’s disease diagnostics.
Blood biomarkers are increasingly being studied not only for whether they can detect disease-related biological signals, but for a more practical question:
Does giving clinicians this information earlier actually change clinical decision-making and patient management?
This distinction is fundamental.
A test that measures more variables is not automatically a better test.
An AI system that answers more questions is not automatically a safer AI system.
The future value of health technology may increasingly depend on three questions:
What did we detect?
How certain are we?
What should happen next?
And sometimes, the most appropriate answer to the third question may be:
“I don’t know. This needs further evaluation.”
In medicine, knowing the limits of one’s knowledge is not necessarily a weakness.
For medical AI, it may be one of the signs that intelligence is beginning to mature.
—
Science & Education:
BI 身体智慧(Body Intelligence)
AI-assisted Research & Illustration:
BI × GPT
Professional Review:
林存默(Thomas Lin)
Professional Community:
ACPN — The Association of Certified Professional Nutritionists
Educational Notice:
This article is intended for professional education and general health information. It does not constitute medical advice, diagnosis, or treatment.
