Start with the failure mode

The most important finding in the ambient-AI literature is not a time saving. It is that only 32% of clinicians given the tool used it in more than half their visits, and that this group got roughly double the EHR-time reduction and triple the documentation reduction of the average adopter. (JAMA, 2026) In one emergency department deployment, the tool was used in 11.2% of eligible encounters, and 62% of attendings never adopted it. (Annals of Emergency Medicine, 2026)

So the design problem is not accuracy. It is abandonment. Everything below is organised around the places a clinician decides, usually without articulating it, that today it is faster to do it themselves.

8:55am — the patient arrives

The patient scans a code at the door or checks in at the desk; both paths land in the same place. Insurance is captured as structured fields — payer, member ID, group, subscriber relationship, Medicare Secondary Payer answers — with the card image retained behind them as evidence. Eligibility runs against the plan being presented today, not the one on file from March.

Intake answers reconcile against the chart rather than accumulating beside it. The screen a staff member sees is a list of differences to confirm, not a form to retype.

Where this fails: when the signature capture silently drops on certain devices and nobody notices until an audit. That is not a hypothetical failure mode; it is a specific bug class worth testing for explicitly, on a real phone, mid-scroll.

9:12am — the visit

The clinician talks through history, exam and plan the way they would explain it to a colleague. The note assembles as they speak. Orders and medications mentioned aloud surface as proposals. The diagnosis and procedure codes take shape from the medical decision making rather than from a separate coding pass afterwards.

Two design decisions matter here more than the model quality.

The draft appears during the visit, not after it. A note that arrives ninety seconds after the patient leaves is a note the clinician reviews in a different mental context, with the encounter already fading. A note taking shape in real time is one they correct as they go.

Every drafted line traces to its source. This is the difference between a thirty-second review and a five-minute one, and thirty seconds is the budget that actually exists. It is also the practical form of FDA's requirement that a clinician be able to independently review the basis of a recommendation.

9:31am — the approval gates

The orders are drafted, not placed. The codes are recommended, not selected. The note is a draft until signed.

This is where most of the argument about clinical AI actually lives, and where the interface choice does more work than the policy. A system that lets a clinician approve twelve proposals with one button has kept a human in the loop on paper and removed them in practice. The automation-bias research is unambiguous that review decays under time pressure, even among experts, even among people who have been trained on the risk.

The measure of an approval gate is not whether it exists. It is whether a tired clinician can get through it without reading.

Our answer is that high-stakes actions get their own explicit confirmation and cannot be batched, and that the recommended E/M level is presented with the decision-making elements behind it rather than as a single number to accept. Slower. Deliberately.

9:34am — the note is done because the visit is

The realistic claim: at consistent use, the documentation for a straightforward visit can be finished at the door, and the clinician's remaining work is editing rather than authoring. The unrealistic claim, which we would rather not make, is that this reclaims the evening — the largest controlled study found no significant change in after-hours EHR time, and the measured average across adopters was closer to a quarter-hour a day than an hour.

What it does reliably change is the shape of the day. Notes that are never in a backlog are never a backlog to face.

Later — the parts that are not the note

Documentation is 2.3 of the 5.8 daily EHR hours in the log data. Chart review, orders and the inbox are the other 3.5. An AI-native EMR that only writes notes is addressing 40% of the problem, which is roughly why so many ambient deployments show smaller gains than expected.

The rest of the day's surface area:

  • Results arrive, get reviewed, and publish to the patient portal or hold — a decision the clinician makes, with a default the practice sets.
  • The claim is a by-product of the visit rather than a reconstruction of it. The visit diagnosis carries into coding directly, which is where a whole class of medical-necessity denials comes from when it doesn't.
  • Denials and authorisations land in a worklist with an owner, and the drafting assistance starts from the note that already exists instead of a blank template.
  • Outbound documents — the referral letter, the fax that a specialist's office still requires in 2026 — get drafted and then wait for a person to send them.

What we would tell a practice evaluating this

Three questions that predict whether you will still be using a system in six months, in descending order of usefulness:

1. How many separate applications are open during a visit? Every one is a place adoption leaks. The products with the worst measured results in head-to-head comparisons are generally not the ones with worse models.

2. What happens on the messiest visit you have? Not the clean follow-up. The one with the interpreter, the family member interrupting, the three unrelated complaints and the medication list the patient describes by colour. Ask to see that one.

3. Can you measure it yourself afterwards? Time to close encounter, after-hours minutes, notes finished at the door. If the system cannot show you those numbers, you will be relying on how it feels — and the evidence says how it feels does not track what the logs show.

What we are not claiming

That this is faster than an experienced physician with a well-built template on a good day. For some clinicians it will not be, and they should keep the template. The case for voice-first documentation is strongest for visits that are conversational, variable and cognitively loaded — which happens to describe most of ambulatory medicine, but not all of it.