2026 / Essay

Who Gets to Call a System Good?

On medical AI, engineering, and deciding what success should mean.

In brief

Imagine a medical chatbot that correctly identifies a possible illness, but the person using it still doesn’t understand whether they should go to the hospital. Did the chatbot succeed? An engineer might point to the correct answer; the person seeking help might say it didn’t solve their problem.

The essay explores that gap between performing well on a test and actually helping someone. It argues that medical AI should be evaluated in the conversations and circumstances where people use it. Patients, clinicians, and engineers should help decide what success means, because accuracy, understandable advice, safety, and access to care all matter. The point is to measure more thoughtfully, rather than assume a high score proves the system is useful.

A medical chatbot can name the right condition and still leave the person using it unsure what to do.

In a February 2026 study published in Nature Medicine, researchers tested three language models on ten medical scenarios with 1,298 UK participants. Given the scenarios directly, the models identified at least one relevant condition in roughly 95% of responses on average. But participants in each model-assisted group identified relevant conditions in fewer than 35% of their final responses. Their choices about what care to seek showed no statistically significant improvement over those of a control group allowed to use their usual information sources. These were simulated situations, not measurements of actual patient outcomes. Still, the gap is striking. Read the study.

The models could produce useful information under one set of conditions. Giving people access to them did not reliably turn that information into better decisions in this study.

Imagine opening a medical chatbot because you are worried about a symptom. You do not arrive with a complete clinical history organized by relevance. You describe what you noticed, leave out something you did not realize mattered, and ask whether you should be concerned. Perhaps you want reassurance. Perhaps you are deciding whether the situation warrants missing work or paying for a visit.

The system has to work with that account, establish what else it needs to know, and communicate something you can use. Naming a plausible condition is one part of that task. Helping you understand what to do next is another.

For someone evaluating the model, identifying a relevant condition might count as success. For the person seeking help, success might mean knowing whether to seek care and how urgently. A clinician might ask whether the exchange missed a warning sign or prompted an unnecessary visit. Each judgment concerns a different part of the same encounter.

Who gets to decide which one makes the system good?

In engineering, choosing a target helps determine what gets built. A test that supplies a complete history and rewards a correct answer encourages us to improve performance under those conditions. But if the intended use involves a conversation with someone who does not know which details matter, gathering information becomes part of the task. So does communicating uncertainty without making the answer impossible to act on.

The distance between those tasks is where some of the most consequential design questions sit. When should the model ask another question? When does a longer answer help, and when does it bury the information someone needs? What should happen when the user's request for reassurance conflicts with the uncertainty in the available information?

These questions matter to model development as well as interface design. They shape which capabilities we train for and which failures our evaluations make visible. A system can improve steadily on the task we specified while leaving another part of its intended use largely untested.

That does not make the narrower test worthless. It makes its scope important.

The healthcare books that prompted this essay helped me put that scope in a larger context. The Health Care Handbook, by Elisabeth Askin and Nathan Moore, provides an account of how American healthcare is organized. Ezekiel Emanuel's Which Country Has the World's Best Health Care? approaches healthcare through comparison across countries. Reading about the organization of care alongside the question of how to judge it raises a problem that also applies to individual technologies: understanding how something operates does not settle what we should ask it to accomplish.

A chatbot might provide a clear recommendation to seek care. Whether that recommendation helps also depends on what care is available, whether the person can afford it, and whether they understand how to obtain it. Those constraints extend beyond the model, but they remain part of the circumstances in which its usefulness is being claimed.

This is not a demand that one tool solve the entire healthcare system. It is a reason to be precise about the benefit we attribute to it. Providing medical information, helping someone choose a course of action, and improving their access to effective care are related achievements. Evidence for one does not automatically establish the others.

Return to the person asking about a symptom. Suppose a chatbot offers a medically reasonable response, explains its uncertainty, and recommends an appropriate next step. That is meaningful progress. We still need to know whether the person understands the recommendation and whether using the tool improves their decision compared with what they would otherwise have done.

We should also ask what the system makes worse. Does it generate so many warnings that people struggle to distinguish levels of urgency? Does it give confident answers when it should seek clarification? Does it work equally well for someone who describes symptoms indirectly or has difficulty reading a long response?

These are questions we can investigate. They do not have to remain expressions of concern.

It would be easy to conclude that benchmarks miss human complexity and should therefore give way to judgment. I do not think that follows. Without measurement, a fluent conversation or a reassuring testimonial can become its own misleading evidence. A tool can feel helpful while leaving the decision unchanged, or making it worse.

The answer also cannot be an unlimited list of things we hope the system will do. Evaluation needs to guide a decision. That requires a primary purpose, a credible comparison, and explicit limits on the harms we are willing to accept while pursuing a benefit.

For a public-facing medical chatbot, one purpose might be helping people choose an appropriate level of care. Identifying relevant conditions would then be supporting evidence, alongside whether users understand the advice and what errors occur. Interviews could reveal why an exchange was confusing; comparative testing could establish whether a revised system actually helps. Following real use, where appropriate, would address questions that a simulated scenario cannot settle.

The difficult part is deciding how to weigh the results. An overly cautious system might avoid some missed emergencies while creating other burdens. A concise answer might be easier to follow while omitting context some users need. Different people may reasonably value those tradeoffs differently.

This is where the question in the title becomes more than a question about testing.

Patients should have a role in defining which burdens and benefits matter. Clinicians bring knowledge of the consequences of errors and the realities of care. Engineers can identify technical constraints and test whether proposed improvements work. The people responsible for delivering services must account for the resources those changes require.

None of these perspectives supplies the whole answer. Including them will not eliminate disagreement. It can, however, make the disagreement visible before one group's priorities become everyone else's definition of success.

I want to build systems whose improvements survive that broader examination. Strong model performance is worth pursuing. So is understanding the conditions under which that performance becomes useful to someone else.

When we call a medical AI system good, we should be able to explain what became better, for whom, and compared with what. We should also know what evidence would make us change our minds.