Stanford pathologists were stuck. Six people had examined a lymph node biopsy, and after roughly 70 cell stains nobody could say what kind of cancer they were looking at. Then a physician testing ChatEHR, Stanford Health Care's homegrown tool for querying and summarizing patient records, asked a question the pathologists had not thought to ask: did the patient have a history of skin lesions? ChatEHR searched the chart and found that a different health system had years earlier diagnosed the patient with sarcomatoid squamous cell carcinoma, a skin cancer known to turn up in exactly this way. "If that doesn't prove the value of ChatEHR, I don't know what does!" the doctor wrote in feedback, in an anecdote that opens Katie Palmer's report this week on how health systems are embracing chatbots built to search patient records.
The anecdote does the tool's marketing for it. A needle in a haystack, found by a machine after trained humans came up empty: that is the promise that has attached itself to generative AI in medicine since the first demos. But the needle is the least representative thing these tools do, and the least important reason to adopt them. The real case for a medical chatbot rests on work that would never make a headline, and the real condition for using one safely is an infrastructure of watching that has nothing to do with miracles.
The records are worst exactly where the tool is most valuable
ChatEHR has a real institutional history, and it did not begin with a diagnostic mystery. Stanford Health Care started building the tool in 2023, led by chief data scientist Nigam Shah and Anurang Revri, on top of its Epic electronic health record system. By mid-January 2026, roughly 1,450 of about 7,000 eligible clinicians were using it, and the system had logged tens of thousands of sessions. Most of those sessions were not mysteries. Clinicians ask whether a patient has an allergy, what the latest cholesterol reading was, what a discharge summary says. The tool summarizes charts for clinicians taking over a case, in the emergency department and for transfer patients arriving with hundreds of pages of records.
Look at the miracle case more closely and it stops looking like a marketing story. The diagnosis that saved the day was buried in another health system's records. External records are the messiest data a clinician ever touches: scanned, inconsistent, filed under different conventions than the home system's. They are also, for exactly that reason, where the highest-value information hides, because no one at the receiving hospital has internalized it. The tool is most valuable precisely where the records are worst. Value and risk concentrate in the same place, and any hospital deploying one of these tools is making a bet that its evaluation apparatus can tell the two apart.
Building it in-house is a bet on owning the error rate
The striking thing about this wave of tools is that the health systems are not buying them. They are writing them. Stanford and Penn Medicine have both built their own chatbots rather than adopting vendor products, and the stated reason is control over failure. Shah has described the vendor landscape for AI tools in health care as "utter madness," with systems juggling too many external tools whose inner workings they cannot see. An in-house tool can be interrogated: its error rates can be measured, its answers traced to sources, its behavior changed in days rather than quarters. Shah's pitch for the homegrown approach is that hospitals should hold the wheel themselves rather than ride as passengers in someone else's vehicle.
Penn's version, Chart Hero, came out of the same frustration from the user side. Srinath Adusumalli, Penn's chief health information officer, has said the existing summarization tools were "very rigid," producing the same fixed summary whether or not it answered the clinician's actual question. Around 70 Penn clinicians are testing Chart Hero, and Adusumalli's stated ambition is to eventually merge chart querying with ambient scribes into a single clinical intelligence platform. Cliniques universitaires Saint-Luc in Brussels has implemented a similar tool, and Duke Health and the Children's Hospital of Philadelphia are developing their own.
The homegrown bet has a price that deserves to be stated plainly. Stanford's build took a team of three full-time employees about four months, and Sutter Health's Ashley Beecy has warned that in-house development is "resource intensive," requiring engineering talent and continuous maintenance that most hospitals do not have. A health system that builds its own chatbot is committing to maintaining software indefinitely, with all the security updates, model upgrades, and regression testing that implies. Owning the steering wheel means owning the crashes. That is the hidden cost of the approach, and it is why the scalability question is open: a well-resourced academic system can staff a small AI team forever, and a community hospital cannot.
The monitoring is the product
Palmer's report carries a subhead that quietly states the whole matter: persistent monitoring is the key to safely deploying AI tools that unearth information buried in health records. This is easy to read as a compliance footnote and wrong to read that way. The monitoring apparatus is the product, and the chatbot is only its input.
Consider what safe use requires. Stanford evaluates ChatEHR against the MedHELM benchmark framework, builds source citations into its answers so a clinician can check where a claim came from, and is developing automations for defined tasks such as assessing transfer eligibility to its Sequoia Hospital and screening for hospice eligibility. Each of those automations is a standing audit: a fixed, repeatable question with a fixed, checkable answer. That is where reliability gets demonstrated, because it is the only setting where being right or wrong can be tallied.
The miracle case demonstrates none of this. A once-in-a-career diagnosis proves that the tool can find a needle, not that it summarizes a medication list correctly ten thousand times a week. The danger is precisely the asymmetry of attention: a clinician will scrutinize a surprising answer, but a routine summary gets accepted at reading speed. The errors that matter for patient care are not the spectacular ones. They are the ordinary ones, and the ordinary ones hide in the work nobody double-checks. Monitoring exists to catch what no single user will.
The clinician's job changes from reading to verifying
There is a second, quieter consequence of these tools that gets less attention than any error rate. A chatbot that summarizes a chart does not just save the clinician time. It changes what the clinician is doing with the time that remains, and the change is not neutral.
Before the tool, the clinician read the record. Reading is active interpretation: the eye moves over the text, and the reader's judgment about which details matter is part of the work. After the tool, the clinician receives a finished summary and verifies it. Verification is a different activity. It is faster, which is the point, but it is also shallower by design, because the summarizer has already decided what belongs in the summary. The decisions the clinician used to make while reading, about what might be relevant to a question not yet asked, are now made partly in advance by a machine that does not know the patient.
This is not an argument against the tools. It is a description of what adopting them means, and it explains why the monitoring question is inseparable from the clinical one. The value of a good summary is measured in minutes saved; the cost of a bad one is measured in the questions that never get asked because the chart was already digested by someone else. The systems that will get this right are the ones that treat summarization as a handoff rather than a substitute: a starting point that points the clinician back into the record at the places that matter, with citations, rather than a finished product to be accepted. That is why Stanford's source citations matter more than the miracle case. A citation is an invitation to verify, and the invitation is the safety mechanism.
The same logic extends to the automations under development. When a tool screens transfer patients for eligibility or flags hospice candidates, it is not answering a clinician's question. It is pre-answering questions the clinician may never have asked, on a schedule no human keeps. The more valuable the automation, the more it operates outside the moment of human attention, and the more the institution's standing audit is the only thing between a wrong pre-answer and an action taken on it.
The easiest thing to demo is the wrong thing to measure
Every health system that adopts one of these tools will be tempted to lead with the miracle. The temptation is understandable; the story is true, and it is the only story about the tool that anyone outside the institution will repeat. But the systems that deploy these chatbots well will be the ones that treat the miracle as a marketing moment and the routine audit as the actual work. The measure of these tools is not whether they occasionally solve a mystery that stumped six pathologists. It is whether they are right, traceably, in the unglamorous thousands of cases a week, and whether the institution watching them catches the failure before it reaches a patient. The diagnostic miracle sells the chatbot. The quiet, standing machinery of verification is what keeps it in clinical use.
Primary sources
- Katie Palmer's STAT article of August 20, 2026, for the Stanford diagnostic case and the treating physician's quoted reaction.
- Palmer's earlier STAT report of January 28, 2026 on Stanford and Penn building their own chatbot tools, for the adoption numbers, Nigam Shah's description of the vendor landscape, Srinath Adusumalli's account of rigid fixed summaries, Ashley Beecy's warning about resource intensity, and the three-person, four-month build at Stanford.
- Stanford Medicine's 2025 news release on ChatEHR for the tool's origins and the MedHELM evaluation approach, and Stanford Health Care's published summary of LLM adoption for the session counts and automation work.
- Healthcare IT News for detail on emergency-department and transfer use.