Hundreds of thousands of American doctors now use clinical AI tools, products from companies like OpenEvidence and UpToDate that are marketed as a safer, medically-grounded alternative to general-purpose chatbots. The pitch is intuitive: a tool built for medicine, trained on clinical literature and guidelines, should be more trustworthy for clinical questions than a consumer model built to do everything. This summer a study put that assumption to a rare head-to-head test, and the result landed hard in the medical-AI world: the specialized clinical tools performed no better, and on several measures worse, than frontier general-purpose models.
It is tempting to read that as a simple verdict, that the general models won. The more careful and more useful reading is almost the opposite. The study is less a ranking of which AI is best than an exposure of how poorly the entire field can measure what it claims to measure. The most important thing it reveals is not that one kind of tool beat another. It is that the tests everyone relies on to make safety claims cannot actually establish which tool is safe.
What the study found, and the caveat in the fine print
The research, from NYU Langone Health and published in Nature Medicine, tested two widely used clinical systems, OpenEvidence and UpToDate Expert AI, against three leading generalist models on a benchmark combining medical-knowledge questions and clinician-alignment tasks. The generalist models came out ahead, and the clinical tools showed deficits in completeness, communication quality, context awareness, and systems-based safety reasoning. The reaction in the health-AI community was intense, because the finding directly contradicts the core marketing claim of an industry that hundreds of thousands of clinicians already rely on.
But sit with one number from the study itself. In the grouped comparison, the generalist models averaged about 94% accuracy and the clinical tools about 89%, and that difference carried a P-value of 0.056, which by the conventional threshold of 0.05 is not statistically significant. In plain terms, the headline result, general beats clinical, sits just on the wrong side of the line researchers normally use to say a difference is real rather than chance. That does not make the study worthless, far from it, but it does mean the decisive-sounding conclusion being repeated is, on the study's own math, closer to a tie than a rout. Anyone citing this as proof that general models are simply better than clinical tools is overreading a borderline result, and the overreading is itself an example of the problem the study should teach.
The benchmark paradox
Here is the deeper issue, and it is the reason the result should not be taken as a clean win for anyone. The benchmarks used to evaluate medical AI may not measure the thing that actually matters in a clinic, and there is direct evidence they do not.
Medical AI benchmarks are typically built from exam-style questions, board-exam items and structured clinical vignettes with known answers. Those are convenient to score, but they resemble the material generalist models are trained on, vast quantities of text including medical exams, more than they resemble the messy, incomplete, high-stakes reasoning of real practice. A model can be excellent at answering a well-formed multiple-choice question and still be dangerous when a real patient presents with ambiguous symptoms, missing history, and no answer key. So a tool "winning" on a benchmark may reflect skill at the test rather than safety in the room.
This is not speculation. A separate 2026 safety benchmark called NOHARM, built from real primary-care-to-specialist consultation cases rather than exam questions, found that severe harm occurred in up to 22% of cases across the models tested, and, crucially, that safety performance was only moderately correlated with the standard knowledge benchmarks, a correlation around 0.61 to 0.64. That moderate correlation is the whole point: it means a model's score on the usual medical benchmarks tells you only part of the story about whether it will harm a patient. A tool can look strong on MedQA-style tests and still carry a meaningful harm rate, because the two things are measuring different capacities. The NOHARM work also found that most errors were harms of omission, leaving out something important rather than stating something false, which is exactly the kind of failure a knowledge-quiz benchmark is poorly designed to catch.
Put those together and the NYU result changes shape. It does not show that generalist models are safe for clinical use. It shows they score slightly higher on a particular kind of test, while a different and arguably more realistic test shows that all of these models, general and clinical alike, harm simulated patients at rates no one should be comfortable with.
The evidence actually cuts both ways
Fairness requires noting that the broader literature does not deliver a single verdict, which is itself the strongest argument against triumphalism in any direction. The NYU study favored generalist models. But other 2026 work points the other way: a benchmark called CSEDB, built on clinical-expert consensus, found that a domain-specific medical model outperformed general-purpose models by a substantial margin, scoring notably higher on safety-critical items like contraindicated medications and dangerous drug interactions. And the NOHARM authors concluded that specialized models can in fact compete with or beat generalist ones on specialized tasks when they are built with genuine in-domain training and evaluated on realistic cases, rather than being thin wrappers around a general model.
So the honest state of the evidence is: it depends on the benchmark, the tasks, and how the specialized tool was actually built. Some studies favor generalist models, some favor specialized ones, and the disagreement among them is not noise to be resolved by picking the most recent headline. It is a signal that the measurement itself is immature, and that confident claims in either direction, whether from clinical-tool marketers asserting superior safety or from generalist-AI boosters declaring victory, are running ahead of what anyone can actually demonstrate.
The clinical tools' fair defense, and the real problem the study exposes
The clinical-AI companies have pushed back on the study, and some of their objections are legitimate. Their tools are often designed to be used in a particular way, integrated into clinical workflows, grounded in cited sources a clinician can check, tuned to defer or surface guidelines rather than to answer confidently in prose. A benchmark that queries them like a chatbot may not capture the value of that design, since a tool built to show its sources and support a clinician's judgment is doing something a raw accuracy score does not measure. That is a reasonable defense, and it should be taken seriously rather than dismissed as spin.
But it also points at the genuine problem the study exposes, which has nothing to do with which model wins. It is that these clinical tools, used by enormous numbers of doctors and sometimes carrying regulatory clearance, have rarely been subjected to independent, quantitative, head-to-head evaluation at all. The marketing claim of superior safety has largely been taken on faith. When independent researchers finally tested it, the claim did not clearly hold up, and the tools' defenders responded that the test did not capture their real strengths, which may be true and which also underscores that no adequate test of those real strengths has been publicly established either. The problem is not that clinical AI lost. It is that, for tools this widely used in a domain this consequential, there is still no agreed, realistic, independent way to know whether any of them is safe enough, and regulatory clearance has functioned as a weak proxy for fitness that this study suggests may not be reliable.
How to read it
The measured way to hold this is to resist the two easy conclusions and sit with the harder one. The easy pro-generalist conclusion, that doctors and patients should just use a frontier chatbot because it beat the clinical tools, is wrong, both because the win was statistically marginal and because the harm-focused evidence shows every model in this space, including the frontier ones, carries real clinical risk. The easy pro-clinical conclusion, that the study is flawed and the specialized tools remain safer as marketed, is also unsupported, because that safety claim was never independently established and did not survive the test that was run.
The conclusion that actually follows is about measurement and humility. Medical AI, in all its forms, has outrun the tools available to evaluate it, and the benchmarks being used to make confident safety claims are demonstrably weak proxies for real-world safety, correlated with it only moderately and blind to the omission errors that cause much of the harm. What the field needs, and what the study most usefully argues for, is transparent, independent, realistic evaluation before these systems are trusted, not benchmark scores repurposed as marketing in either direction. For the clinicians and patients relying on these tools right now, the practical takeaway is neither to switch to a general chatbot nor to trust the clinical label, but to treat all of them as capable, useful, and not yet proven safe, and to keep a qualified human firmly in the loop. The study did not tell us which AI to trust. It told us we do not yet have a trustworthy way to decide, and in a field moving this fast, recognizing that is more valuable than a winner.
Primary sources
- STAT for the framing that hundreds of thousands of US doctors use clinical LLMs from companies including OpenEvidence, Doximity, and UpToDate, that these are pitched as an antidote to generalist-model hallucination, that the NYU Langone study published in Nature Medicine found clinical AI performed worse than general models, and that the paper triggered unusually strong reactions in the health-AI community.
- The NYU Langone preprint and ResearchGate/arXiv postings for the 1,000-item benchmark combining MedQA and HealthBench, the comparison of OpenEvidence and UpToDate Expert AI against three leading generalist models, the grouped accuracy of about 94.1% for generalist versus 89.0% for clinical tools at P=0.056, and the identified deficits in completeness, communication, context awareness, and systems-based safety reasoning.
- The NOHARM benchmark paper for the 100 real consultation cases across 10 specialties, severe harm in up to 22.2% of cases, harms of omission accounting for about 77% of errors, the moderate 0.61-0.64 correlation between safety and existing benchmarks, and the finding that well-built specialized models can compete with generalist ones.
- npj Digital Medicine's CSEDB study for the finding that a domain-specific medical model outperformed general-purpose models, particularly on safety-critical items.
- Clinical Trial Vanguard for the regulatory framing that FDA-cleared clinical tools underperformed general-purpose models and that clearance may be an unreliable proxy for fitness-for-purpose.