More than 40 million Americans ask ChatGPT a health question every day, mostly after clinic hours, according to a STAT First Opinion essay published this week by Arya Rao and Marc Succi, physicians at Harvard Medical School and Mass General Brigham. Their subject is the shadow medical system being assembled from consumer parts. Oura sells a 50-biomarker blood panel drawn at Quest Diagnostics locations for $99. Ro and Hims write prescriptions for weight loss or anxiety after an asynchronous intake. Function Health, valued at $2.5 billion in November, lets members order 160 lab tests a year and authorize ChatGPT to read the results. Doctronic, which calls itself "the world's #1 AI doctor," says it has run 24 million consultations and now writes AI-generated prescription refills in Utah.

The rush is easy to explain. Health care is close to one-fifth of the American economy, and every frontier AI company is moving into it. What is harder to explain is what the race is being scored on. The claims that drive this market, the "AI doctor outperforms ChatGPT" headlines behind the product launches, come from benchmarks. The systems are ranked on tests, and the test is not the practice.

The scoreboard runs on exam questions

The field's most influential scoreboard is MedQA, a dataset introduced in 2020 by researchers at MIT that assembles roughly 12,700 multiple-choice questions from professional medical licensing exams, USMLE material included (the MedQA paper). The questions were drawn from public test-prep sites rather than from hospitals or clinics. For years the milestone was simply passing: Google's Med-PaLM was the first model to clear the passing threshold on USMLE-style questions, and Med-PaLM 2 later pushed past 86 percent. Those results were treated as scientific events.

The ceiling has since become routine. A study published in Nature Medicine this June found general-purpose frontier models outscoring FDA-cleared clinical AI tools on MedQA and other medical benchmarks: Gemini 3.1 Pro scored 97.4 percent, GPT-5.2 scored 94.2 percent, while the cleared tools OpenEvidence and UpToDate Expert AI stayed under 90 percent (the Nature Medicine study). A Scientific Reports comparison this year had DeepSeek beating ChatGPT on both the USMLE and China's national medical licensing exam, 92.6 to 90.3 percent and 86.8 to 79.4 percent respectively. Version by version, the race is real and the numbers move.

Those numbers are also the marketing. They are the ones quoted in product pages, cited in funding announcements, and repeated in the coverage that makes a consumer think a system has reached doctor level. And the numbers are accurate as far as they go. What they do not do is measure the thing a doctor actually does.

The exam ends where the visit begins

Rao and Succi know this gap from the inside. Their own study, published this spring in JAMA Network Open, is the largest evaluation to date of frontier models across the full arc of clinical reasoning (the JAMA Network Open study). They tested 21 off-the-shelf models on 29 clinical vignettes and scored them across five reasoning domains, from differential diagnosis through management, using a composite they call the PrIME-LLM score.

The results split cleanly along the format of the question. Given a complete case with labs, imaging, and organized data, the models named the correct final diagnosis more than 90 percent of the time. Given only what a clinician gathers in the opening minutes of a visit, they failed to produce a comprehensive differential more than 80 percent of the time. An invited commentary in the same journal notes that even the strongest reasoning models struggled to generate appropriate differentials, and that a model can land the correct final diagnosis while failing to construct a coherent differential, which suggests recall of a known answer rather than reasoning toward one (the JAMA commentary).

The differential is where care actually starts. It decides which tests get ordered, what the physician probes, what gets ruled out first. A wrong or empty differential delays care, drives unnecessary testing, and sends patients home with the reassuring final answer to a question no one asked. The authors' own warning is that the strong final-answer scores create a misleading sense of safety, because the system is being scored at the moment care ends rather than at the moment it begins. The exam hands the model a complete picture and five clean options. A visit hands a clinician fragments, noise, and a patient who does not know what to say.

The clinic gets resold as a product

The consumer layer is now built around this same confusion between taking a test and receiving care. Quest Diagnostics launched an AI companion in its MyQuest app this spring that reads up to five years of a user's lab results and answers questions about them, powered by Gemini. Hims and Hers added a Labs AI tool that interprets lipid panels and metabolic markers. Oura's blood panel wraps a lab draw in app-based insights and AI summaries. Function Health goes furthest, coupling 160 tests a year with a full-body MRI and permission for ChatGPT to read everything.

Each of these products delivers a result: a number, a trend, a flagged marker, a summary. None of them delivers the interpretation a physician would give, because none of them carries the physician's responsibility. That distinction is precisely what has eroded in the underlying models. A longitudinal study of 15 AI models across 500 health questions and 1,500 medical images found that medical disclaimers in chatbot health answers fell from 26.3 percent in 2022 to 0.97 percent by 2025, and the researchers found a striking inverse relationship: the more accurate a model became, the fewer disclaimers it offered (the disclaimers study). The better the exam score, the more confident the tone, with nothing left on screen to say otherwise.

Fairness requires noting that the picture is not uniform. A separate evaluation of newer models on real patient queries found that urgent questions drew more explicit disclaimers and referrals than casual ones, with 97 percent of responses pointing a user to a medical professional (the urgency study). And the improvement across three years is genuine: reasoning-tuned models scored higher, and most models handled imaging better than text alone. The optimistic case is real progress. The pessimistic case is that the progress is being sold as a substitute for the piece that did not improve, which is the piece the patient actually needs.

Even the strongest result is a company's own test

The sharpest example of the benchmark-as-battleground is Doctronic. Under a 2024 Utah regulatory sandbox that waives parts of scope-of-practice law, the company runs the country's first pilot in which an AI chatbot renews prescriptions for chronic conditions, limited to about 190 common medications and reviewed by humans in the program's early phase. The company expects to drop the human review over time, and the pilot's oversight board is made up of AI specialists, none of them physicians. The company's public case rests on a preprint comparing its agent against board-certified clinicians across 500 virtual urgent care encounters, reporting 99.2 percent treatment-plan consistency with the clinicians' plans, with the comparison judged by another large language model (the Doctronic preprint).

The study is not peer-reviewed, it is authored by the company, and it is its only published study. Critics have catalogued the limitations: the encounters were low-acuity telehealth visits, the judge was itself an AI, and the clinicians saw the AI's notes before making their own plans, which can anchor their judgment. The American Medical Association has warned that "prescription renewals aren't routine checkboxes," and Utah's medical board, which says it was not consulted before the launch, has called for the program to be halted. One physician put the threshold question plainly: giving something that is not human a medical license crosses a line, and the company's own sales pitch, "the world's #1 AI doctor," shows how the scoreboard and the pitch have merged.

None of this requires dismissing the company's case. Its defenders argue for graduated autonomy with clinicians at the center, and the four divergent cases in the preprint involved broader workups rather than immediate harm. The point is narrower: Doctronic's strongest evidence is its own benchmark, judged by its own standard, and the state is relying on it. An IBM training manual from 1979 said a computer must never make a management decision because it cannot be held accountable for one. That sentence is older than every company in this story, and it has not aged.

Whoever controls the test controls the market

This analysis takes no position on whether consumer health AI should be regulated, licensed, or left alone. The observation is more mechanical. When a market is scored on a benchmark, the winner is whoever controls the benchmark: who writes the questions, who owns the judge, who publishes the score. Today that control sits with the vendors. The consumer cannot distinguish "passed the exam" from "handles care," because the two numbers look identical and only one of them is published.

The fix is not to abandon benchmarks. It is to move the scoreboard to where care begins: partial information, follow-up questions, the willingness to say "I do not know," the discipline to order the right test and the right time, and the judgment to hand the patient to a human at the moment the stakes rise. And when a wrong answer causes harm, the one who answers for it is still the physician, not the company that published the score. Rao and Succi's own prescription is a supervised system, working between visits and behind the scenes, not one at the front of the examination room. Scoring that takes different evaluations, and several exist: blinded clinician review of real patient queries, dialogue-based assessments, and structured observation of early differentials.

Until those scores are the ones quoted in headlines, the number the consumer sees will be a number on a test. And the patient asking a chatbot at 11 p.m. whether to worry is not taking a test. They are at the start of a visit, with fragments, noise, and no answer options. The race is real, and the systems are genuinely improving. But the contest is being scored on the wrong test, and the scoreboard belongs to the people selling the score.

Primary sources

  1. STAT's First Opinion essay "AI has created a shadow medical system," Arya Rao and Marc Succi, August 19, 2026, for the usage figures, the consumer health features at Oura, Quest, Ro, Hims, and Function Health, the Doctronic self-description, and the summary of the authors' JAMA study.
  2. "Large Language Model Performance and Clinical Reasoning Tasks," Rao, Esmail, Lee, et al., JAMA Network Open 2026;9(3):e264003, for the PrIME-LLM scoring, the final-diagnosis results, and the early-differential failure rates, with the accompanying commentary by Mickael Tordjman and Xueyan Mei providing the interpretation of the reasoning gap.
  3. Sharma, Alaa, and Daneshjou's "A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models" for the disclaimer-decline statistics, and a Journal of Medical Internet Research study for the countervailing finding on urgency-based disclaimers.
  4. Jin et al.'s "What Disease Does This Patient Have?" for the origins of the MedQA benchmark, the June 2026 Nature Medicine study for the comparative benchmark scores, and the Scientific Reports comparison of DeepSeek and ChatGPT for the exam-race figures.
  5. The medRxiv preprint "Toward the Autonomous AI Doctor" for the 99.2 percent treatment-plan consistency result, and Associated Press reporting for the details of the Utah sandbox, the medical board's response, and the American Medical Association's warning.