What the rules actually say now
A surprising amount of published advice about office-visit coding describes rules that were replaced years ago. The current framework, for 99202–99215:
- History and physical exam do not determine the level. Since 2021, CPT requires "a medically appropriate history and/or physical examination, when performed," and states explicitly that its extent is not an element in code selection. The treating clinician decides what is medically appropriate.
- The level comes from medical decision making or total time, the clinician's choice, per encounter.
- MDM has three elements — the number and complexity of problems addressed, the amount and complexity of data reviewed and analysed, and the risk of complications from patient management. Two of the three must be met or exceeded.
- Total time counts only the physician's or QHP's own time on the date of the encounter, face-to-face and not, including chart prep, ordering and documentation. It excludes clinical staff and scribe time.
- Since 2024, the time descriptors are minimum thresholds that must be met or exceeded, not ranges: 99213 at 20 minutes, 99214 at 30, 99204 at 45, 99205 at 60.
Note what changed in 2023 and what didn't: the 2023 revisions extended the office-visit framework to hospital, observation, consultation, nursing facility and home services. Office visits themselves were not re-rewritten.
Office visits are the category where coding, not charting, is the problem
CMS's Comprehensive Error Rate Testing programme reviews a sample of Medicare fee-for-service claims each year. For established-patient office visits in the most recent published cycle, 65.2% of improper payments traced to incorrect coding, against 18.4% with no documentation and 16.4% insufficiently documented. Medical necessity accounted for 0.0%. (CMS, 2026)
That is unusual. Across Medicare as a whole, insufficient documentation is the single largest root cause of improper payments. For office visits it is a distant second. The service is documented; the level is wrong.
Direction matters here, and so does a caveat most write-ups omit. In the CERT sample of established office visits, 120 claims were coded above what the documentation supported and 7 below — roughly 94% overcoding among detected errors. But CERT is an improper payment instrument. It is built to find money Medicare should not have paid. It is structurally poor at finding revenue a practice never billed for. Anyone quoting CERT as proof that undercoding is rare has misread what the programme measures.
A number to be careful with
The widely circulated figure that 42% of Medicare E/M claims are miscoded — 26% up, 14.5% down — comes from an OIG review of 2010 claims, a decade before the office-visit rules were rewritten. (OIG, OEI-04-10-00181) It is still cited as current. It isn't.
The part nobody says out loud: the answer is genuinely ambiguous
In a 2025 study, an independent expert coder reviewed 117 real de-identified outpatient encounters that had already been coded by a professional coder. The expert disagreed with the original code on 56% of encounters. The authors note that determining the MDM level is the most complex and error-prone step. (Nassar et al., EMNLP 2025)
If two qualified humans reach different levels on the same chart more often than they agree, then "coding accuracy" is not purely a measure of physician error. Part of it is irreducible judgment.
This reframes the whole problem. A tool that promises to eliminate coding error is promising to resolve an ambiguity that professional coders have not resolved among themselves.
What AI is measurably bad at here
The evidence on language models and medical codes is not close.
A benchmark published in NEJM AI tested GPT-4, GPT-3.5, Gemini Pro and Llama-2-70b on reproducing codes from clinical descriptions, drawing on more than 27,000 unique diagnosis and procedure codes from a year of routine care. Every model scored below 50% across ICD-9-CM, ICD-10-CM and CPT. (Soroush et al., 2024)
E/M coding specifically is hard enough that the research literature tends to report improvement over a baseline rather than absolute accuracy. In the study above, the best purpose-built system scored 33.7 percentage points higher than a single-prompt GPT-4o baseline and 36.9 points higher than a commercial tool — on a 99-encounter test set, against a ground truth that expert coders themselves disagree about more than half the time. The paper publishes no absolute accuracy figure for any system, which is itself informative.
Set that against the marketing, where autonomous-coding accuracy above 95% is commonly advertised and, as far as we can find, never independently validated. The published benchmarks point the other way. We have no interest in joining that particular arms race.
So what is AI genuinely good for in coding?
Assembling the evidence, not choosing the answer. The useful division of labour looks like this:
The model organises the medical decision making
Problems addressed and their complexity, data reviewed, risk of management — pulled from what was actually said and done in the visit, laid out against the two-of-three rule, with each element traceable to its source in the encounter. This is a summarisation problem, which language models are good at, rather than a classification problem, which they are not.
The model shows the level that follows, and the one next to it
Presenting a recommended level with the reasoning is useful. Presenting it as the answer is not. Where the encounter sits near a boundary — and a great many do — the software should say so rather than pick a side silently.
The clinician selects the code
This is not merely a compliance posture, though it is that too. It is the only defensible arrangement when the ground truth is contested. The clinician was in the room; the model was listening to it.
Time is documented, not inferred
If a level is being supported by total time, the time attested must be the clinician's own time on that date, and the record needs to show it. A tool that quietly counts scribe time or staff time into a total is building an audit finding.
The uncomfortable finding about coding and ambient AI
A UCSF study of 1.2 million encounters found ambient documentation associated with a small but real revenue increase — about 1.81 additional RVUs per physician per week, with no increase in claim denials. The authors were candid that they could not distinguish expanded services from improved coding accuracy from inappropriate upcoding. (JAMA Network Open, 2026) Any vendor selling documentation AI on a revenue-lift argument owes you that sentence.
How we do it
ChartVoyant proposes an E/M level from the medical decision making it can evidence, shows the elements it relied on, and asks the clinician to choose. New-patient and established-patient paths are handled separately, because they are separate code families with separate thresholds. The model never selects the code, and it never silently moves a level between draft and signature.
That is a slower design than autonomous coding. Given what the published accuracy numbers look like, we think slower is the correct answer.