1 What “accurate” has referred to
When a clinical software vendor describes its language model as accurate, the claim usually resolves to a score on a multiple-choice examination. That is not a rhetorical complaint. It is the measured shape of the literature.
A systematic review published in JAMA in 2025 assessed 519 studies of large language models in health care. Only 5% used real patient care data; 44.5% used medical examination questions. Accuracy was the primary metric in 95.4% of studies, and only 15.8% assessed fairness, bias, or toxicity. The task distribution is more telling: question answering 84.2%, summarization 8.9%, dialogue 3.3%, billing codes and prescriptions 0.2% each.1 Those last two describe much of what an ambulatory practice delegates to software.
This paper makes three claims. The headline figures in circulation descend from a narrow lineage of examination-derived benchmarks whose provenance is routinely misstated. The properties that would make a benchmark score interpretable, contamination control chief among them, are missing from most medical benchmarks. And the few studies that measured clinical work rather than examination performance yield a less flattering, more useful picture. It is a review of others’ published research and contains no new data.
2 Where the famous numbers come from
Three papers supply most of the accuracy figures repeated in sales material. Each is misquoted in a specific way.
The Kung paper is not “ChatGPT passed the USMLE”
Kung et al. tested ChatGPT on 350 USMLE items drawn from a starting set of 376.2 These were practice materials rather than the administered examination, and the model was GPT-3.5. Censoring indeterminate responses gave Step 1 75.0%, Step 2CK 61.5%, and Step 3 68.8%; without censoring the same figures were 45.4%, 54.1%, and 61.5%. The paper described this as performance at or near the passing threshold, which is around 60% and varies by Step and by year.
The popular claim therefore fails three ways: the items were practice items, the uncensored figures for two of three Steps fall below the threshold, and the study’s own language stops short of the claim.
The twenty-point margin belongs to a vendor preprint
The claim usually paired with Kung, that the model exceeded the passing score by more than 20 points, comes from Nori et al., who evaluated GPT-4 on USMLE practice materials and MultiMedQA.3 That work is an arXiv preprint from March 2023, not peer reviewed, authored by the research laboratory of the company that produced the model. The figure carries the weight of a vendor self-report, and attributing it to Kung merges two papers testing two different models.
Med-PaLM 2 did not outperform physicians
Singhal et al. introduced MultiMedQA and reported Flan-PaLM at 67.6% on MedQA, more than 17 points above the prior state of the art; the paper’s own framing was that this exposed the limits of the benchmarks rather than clinical readiness.4 The follow-up reported Med-PaLM 2 at 86.5% on MedQA, 72.3% on MedMCQA, and 75.0% on PubMedQA, rising to 81.8% with self-consistency.5
The “outperforms physicians” claim attaches to a different measurement in the same paper. Across 1,066 MultiMedQA questions, physician raters compared long-form answers pairwise and preferred Med-PaLM 2 on 8 of 9 axes; on better reflecting medical consensus it was preferred 72.9% of the time. A separate bedside-consultation set of 20 questions had specialists preferring it to generalist physician answers 65% of the time.5 That is a rater preference over text: no patient, no encounter, no outcome. The 86.5% measures multiple-choice accuracy, the preference result measures prose, and the two are routinely compressed into one sentence that neither supports. Ayers et al. is the precedent: raters preferred chatbot answers in 78.6% of 585 evaluations, but the physician comparators were unpaid volunteers on a public forum averaging 52 words against the chatbot’s 211.6
3 What a benchmark score does not control for
MedCheck audited 56 medical LLM benchmarks against a lifecycle rubric and found an evaluation infrastructure largely without controls. 88% (49 of 56) had no contamination-risk mitigation of any kind, 89% could not test robustness, 91% did not evaluate uncertainty awareness, half did not align with ICD or SNOMED CT terminology, only 38% used authentic real-world scenarios, and 80% had no maintenance plan.7 That audit is itself a non-peer-reviewed preprint.
Why contamination makes a score an upper bound
A benchmark score is meant to estimate performance on unseen items. If the items were in the training data, the score measures retrieval rather than generalization, and the measured value is at least as high as the true one. The direction of the error is known; the magnitude is not, because nobody measured it. A contaminated benchmark yields an upper bound of unknown tightness, which is a weaker object than an estimate.
Examination items are the worst case: widely reproduced online, stable across years, likely to be in a large web crawl. Set beside the finding that 44.5% of the literature runs on examination questions while 5% uses real patient data,1 the field’s central quantity is an upper bound of unknown tightness on a task nobody deploys.
Benchmarks aimed at the actual work
MedHELM is the most developed attempt to cover clinical work: a clinician-validated taxonomy of 5 categories, 22 subcategories, and 121 tasks, assembled from 37 benchmarks and applied to 9 frontier models. Its argument is that licensing-examination benchmarks do not represent real clinical work.8 The 15.8% fairness figure carries a cost accuracy metrics cannot detect. Omiye et al. put nine questions to four models across five runs each and found all four reproduced debunked race-based claims, including race-adjusted eGFR justified by false assertions about muscle mass in Black patients; the authors concluded the models were not ready for clinical use.9 No MedQA item asks any of that.
4 What happens when the record is real
Hager et al. evaluated models on 2,400 real patient cases from MIMIC-IV across four abdominal pathologies.10 Hospitalists achieved 87.5% to 92.5% diagnostic accuracy; the open models achieved 58.8% (Llama 2 Chat), 67.8% (OASST), and 65.1% (WizardLM), a gap of 16 to 25 percentage points. Accuracy fell further when a model had to gather information autonomously rather than receive it, from 58.8% to 45.5% in one case. Laboratory-value interpretation ranged from 24.1% to 77.2% depending on model and direction, guideline adherence was inconsistent, and the models hallucinated nonexistent tools every two to five patients.
The autonomous-collection result transfers directly to purchasing decisions. A demonstration presents a curated case: relevant history assembled, pertinent labs surfaced, question well posed. Live use requires the system to decide what to look at, which is the step at which accuracy fell. A demonstration that hands the model its context measures a task the deployed system will not face.
Tool hallucination is a separate failure mode: not a wrong answer but a confident invocation of a capability that does not exist. Inside an EMR that is an order or an interface the model believes is available, so the error surfaces as an action rather than as a sentence.
Limitations
This study evaluated open-weight models available at the time, on four abdominal pathologies, retrospectively, against structured MIMIC-IV records. It establishes that the gap is measurable and that it widens when the model must gather its own information. It does not establish the size of that gap for current proprietary models, for other presentations, or in prospective use. No study reviewed here repeats the design on current frontier models, so the present gap is unknown rather than small.
5 Hallucination, and the denominator problem
“Large language models hallucinate X% of the time” is not a claim that can be evaluated, because the published rates use at least four incompatible denominators. Table 1 sets them side by side for the sole purpose of showing that they cannot be placed side by side.
| Study | Denominator | Rate | What limits the reading |
|---|---|---|---|
| Asgari 202511 | Sentences in generated summaries (12,999 annotated) | 1.47%, of which 44% graded major | Scales with note length; major errors clustered in the Plan section (21%) |
| Palm 202512 | Notes containing at least one hallucination | 31% of ambient notes vs 20% of physician notes (p=0.01) | Note-level presence, not error density; 97 encounters, 194 notes |
| Chelli 202413 | Generated bibliographic references | GPT-3.5 39.6%, GPT-4 28.6%, Bard 91.4% | Reference generation is a distinct task from note generation |
| Bhattacharyya 202314 | Generated references (115 across 30 short papers) | 47% entirely fabricated; 46% authentic but inaccurate | Two error classes with different consequences, collapsed by casual citation |
| Omar 202515 | Adversarial vignettes seeded with one fabricated element | 50–53.3% to 80–82.7% by model; 66% overall | Deliberately adversarial input; an upper bound on elaboration, not a base rate |
The per-sentence figure comes from 12,999 clinician-annotated sentences across 450 note and transcript pairs, with the major hallucinations concentrated in the Plan section. Omission was more than twice as common, at 3.45% of 49,590 annotated items.11 Omission appears in no vendor hallucination statistic and is harder to catch on review, since nothing incorrect is present to notice.
The note-level figure comes from a blinded comparison of 194 notes across 388 paired specialist reviews. Reviewers preferred the ambient note 47% of the time against 39% for the physician draft, and rated overall quality at near parity, while those same ambient notes carried more hallucinations.12 Preference and error rate moved in opposite directions within one study, a warning about evaluating documentation tools by reviewer preference.
The adversarial result matters most for design. Omar et al. seeded each of 300 physician-validated vignettes with a single fabricated element, a fake laboratory value, sign, or condition, and measured whether the model elaborated on it. Elaboration ran from 50–53.3% for GPT-4o to 80–82.7% for a distilled DeepSeek-R1, 66% overall, and a mitigation prompt cut the overall rate to 44%. Setting temperature to zero produced no significant improvement.15 Determinism is not accuracy. A temperature-zero model returns the same output every time, including the same fabrication.
6 The input layer: speech recognition
Every claim about a voice-driven tool rests on a transcript, which has its own error rate, usually omitted from the tool’s accuracy figure. Zhou et al. remains the reference measurement: 217 clinical documents comprising 83 office notes, 75 discharge summaries, and 59 operative notes. Errors per 100 words were 7.4 in raw speech-recognition output, 0.4 after professional transcriptionist editing, and 0.3 in the physician-signed note. Quoted alone, that sequence reads as a review pipeline working.
The second series in the same study is the one that matters. The proportion of remaining errors that were clinically significant was 5.7% at recognizer output, 8.9% after the transcriptionist, and 6.4% in the signed note.16 Human review cut error volume roughly twentyfold and did not reduce the share of surviving errors that carried clinical consequence. Review removes the errors that are easy to see; a plausible wrong laterality survives a read-through that catches a garbled word. A pipeline reporting a post-review error rate reports a quantity that improves while the risk does not.
Goss et al. found 1.3 errors per note in 100 emergency department dictated notes, up to 15% of errors judged critical, and clinicians perceiving error rates as lower than measured; those figures reach this paper at second hand and should be checked against the original.17
Whisper is often cited here, and usually incorrectly. Koenecke et al. found it produced entirely hallucinated phrases or sentences in roughly 1% of transcriptions, text with no counterpart in the audio, and that 38% of those hallucinations contained explicit harms, including perpetuating violence and implying false authority; fabrication was disproportionately triggered by longer non-vocal pauses.18 That study used aphasia speech corpora and control speech, not medical dictation. It does not establish an error rate for clinical transcription, and the circulated accounts of Whisper fabricating in hospital transcription derive from journalism rather than from this paper. What it establishes is a mechanism and a disparate impact: any speaker who pauses is at elevated risk, patients with aphasia among them. In ambient documentation the disfluent speaker is frequently the patient.
7 Trials that measured outcomes rather than scores
A small number of randomized trials measured what happened when clinicians used these systems. They do not point in one direction.
| Trial | Design and participants | Primary comparison | Result |
|---|---|---|---|
| Goh 2024, diagnostic reasoning19 | Single-blind RCT; 50 physicians, 6 vignettes, 244 cases | LLM plus conventional resources vs conventional resources | 76% vs 74% median; adjusted difference 2 points (95% CI −4 to 8; P = .60). LLM alone 92% |
| Goh 2024, management reasoning20 | RCT, preprint; 92 attendings and residents | GPT-4 access vs no access | +6.5 percentage points (95% CI 2.7–10.2, P<0.001); +119.3 s per case |
| Everett 202621 | RCT; 70 clinicians, 254 cases | Three workflow positions vs conventional resources | Baseline 75%; second opinion 82%; first opinion 85%; AI alone 87% |
| Agweyu 202622 | Pragmatic cluster RCT; 16 facilities, 103 clinical officers, 9,347 patients analyzed | Generative CDS vs usual care, real patients | Treatment failure at 14 days 2.2% vs 2.0%; adjusted OR 0.77 (95% CI 0.55–1.08, P = 0.13) |
The diagnostic-reasoning trial is routinely reported as evidence that AI beats doctors. The accurate reading is narrower and more useful: physicians using the model did no better than physicians without it, and no faster (519 seconds against 565, P = .20), while the model alone did better than both arms.19 That is an integration failure, not a capability claim, and it is invisible to every benchmark described above. A benchmark measures a model; this trial measured a model plus a clinician plus an interface, and the combination underperformed its best component.
The opposite generalization is equally unsupported. Goh’s management-reasoning trial found management decisions up 6.1% (P=0.001) and diagnostic decisions up 12.1% (P=0.009), with augmented physicians not differing from GPT-4 alone (−0.9%, P=0.8). It is a medRxiv preprint, and its publication status should be confirmed before it is cited as peer-reviewed.20
Everett et al. suggests the operative variable is workflow position rather than model quality. Against a 75% baseline, AI as a second opinion reached 82% (+6.8%, 95% CI 4.0–9.6%) and AI as a first opinion 85% (+9.9%, 95% CI 4.7–15%), while AI alone reached 87%.21 The same class of model produced different collaborative results depending on when the clinician saw its output: a property of the product, not of the model.
The one trial with real patients and a clinical primary outcome is the least flattering and the most relevant. Its secondary outcomes showed improved documented diagnostic quality and treatment planning, with no differences in antibiotic or antimalarial prescribing, and its 33 serious adverse events were adjudicated as unrelated to the intervention.22 Across all four trials, documentation quality and rated reasoning move more readily than patient outcomes.
Limitations
Three of these four trials used vignettes rather than patients, and their outcome is a reasoning score rated by humans rather than a clinical event. The fourth used real patients, but its primary outcome was treatment failure at 14 days with a low event rate, 2.0% in control, so a null result there stays compatible with small true effects in both directions. None of the four measured billing accuracy, prescription safety, or documentation burden, leaving the highest-volume administrative applications without randomized evidence.
8 Reporting standards, and what they would and would not fix
Four reporting guidelines now cover this territory, and they are frameworks rather than evidence. TRIPOD-LLM extends TRIPOD+AI to language models with a modular checklist covering development, tuning, prompting, and evaluation.23 DECIDE-AI covers live clinical evaluation between offline validation and large comparative trials, the stage most relevant to software already shipping AI features, with 27 reporting items.24 CONSORT-AI added 14 items to CONSORT 2010 through a two-round Delphi and a 34-participant pilot.25 SPIRIT-AI added 15 items to SPIRIT 2013 through the same process, covering trial protocols rather than trial reports.26
Applied to the studies in section 2, these checklists would have forced disclosure of the items used, the model version and evaluation date, whether contamination was assessed, how indeterminate responses were handled, and who funded and authored the evaluation. That is most of the correction work this paper has done by hand.
What they do not do is measure anything. A reporting guideline improves the description of a study, not the study, and adherence is voluntary. A perfectly reported evaluation of a contaminated benchmark is still an evaluation of a contaminated benchmark.
The assurance proposals aim at that gap. Shah et al. proposed a nationwide network of health AI assurance laboratories issuing performance reports usable across populations and settings.27 The Coalition for Health AI published an Assurance Standards Guide setting out a six-stage lifecycle from problem definition through ethical design, engineering, deployment and monitoring; it is an industry consensus standard, not peer-reviewed research.28 Both remain proposals. Neither has produced published evidence that a model reviewed under one performs better in deployment.
9 What follows
The answerable question is not whether a clinical language model is accurate. It is: on what task, measured how, against which comparator, and when. The literature reviewed here supports five properties that separate an evaluation worth reading from a number worth ignoring.
- Task match. The evaluation covers the work the software will perform rather than a proxy for it. Coding and prescribing carry the thinnest published evidence and are among the most commonly automated tasks.1
- Data provenance. Real records rather than curated vignettes, and an explicit statement of whether the model was handed its context or had to gather it.10
- A stated denominator. Every error rate names its unit: per sentence, per note, per reference, per case. A figure quoted without its unit conveys nothing.11,12,13,14,15
- Contamination handling. Either evidence that the evaluation items could not have been in the training data, or a plain statement that this was not assessed, in which case the score is an upper bound rather than an estimate.7
- An outcome, or an honest absence of one. Reasoning scores and preference ratings measure real things, none of which is a patient outcome. A null outcome is information about the intervention, not a defect in the study.19,22
None of this requires evidence that does not yet exist. The reporting frameworks are published, benchmarks aimed at clinical content exist, and randomized trials with clinical endpoints have been run. What is missing is coverage: the field has measured examination performance extensively and the clinical and administrative work these systems are sold to do barely at all. Until that changes, an accuracy claim should be read as a statement about a benchmark, and a demonstration as a statement about a curated case. Neither is a statement about a practice.