The study everyone quotes

In 2016, a time-motion study in Annals of Internal Medicine put observers in the room with 57 physicians across four specialties for 430 hours. The finding that escaped into general circulation: physicians spent 49.2% of the office day on EHR and desk work and 27.0% on direct clinical face time. Even inside the exam room, with the patient sitting right there, 37% of the time went to the computer. (Sinsky et al., 2016)

A year later, a study using EHR event logs rather than human observers found primary care physicians spending 5.9 hours a day in the record — 4.5 during clinic and 1.4 after it — implying an 11.4-hour workday. (Arndt et al., 2017)

Those numbers are now a decade old, and they get cited as though nothing has changed. It's worth checking whether that's true.

What the logs say now

It is true, roughly. The largest EHR log study yet conducted — 200,081 physicians across 396 organizations — found 5.8 hours of active EHR time for every 8 hours of scheduled patient time: 2.3 hours documenting, 1.1 reviewing charts, 0.8 on orders, 0.8 on the inbox. Only 57.8% of that happened during scheduled hours. About a fifth landed outside scheduled hours on clinic days, and another fifth on days with no clinic at all. (JGIM, 2024)

The same research group produced a metric that is more useful to a practice owner than any percentage: PSH40, the number of scheduled patient hours that add up to a 40-hour week once the EHR work is counted. The median is 33.2 hours. Schedule a physician for 36 and you have scheduled a 43-hour week, most of the overage invisible. (JAMIA, 2025)

Every hour you put on the schedule quietly brings about eighteen minutes of unscheduled work with it. That ratio, not the visit length, is what sets the length of the day.

Pajama time is the part that hasn't moved

The AMA's Organizational Biopsy — nearly 18,000 physician responses in 2024 — found 22.5% of physicians spending more than eight hours a week in the EHR outside 5:30pm to 7am, up from 20.9% the year before. Over the same period the average total work week fell, from 59 hours to 57.8. The AMA's own framing was that doctors are working fewer hours but the EHR still follows them home. (AMA, 2024)

This matters because after-hours work is the burden with the clearest link to attrition, and attrition is the expensive part. A modelling study in Annals put the national cost of burnout-attributable physician turnover and reduced clinical hours at $4.6 billion a year, or roughly $7,600 per employed physician. (Han et al., 2019) In a five-physician practice, one departure resets the year.

What ambient AI has actually been measured to do

Here the evidence is genuinely encouraging and genuinely more modest than the marketing.

The strongest results are on how the work feels. A six-system study found burnout falling from 51.9% to 38.8% within 30 days of ambient documentation going live, with cognitive task load down and focused attention up. (JAMA Network Open, 2025) At Mass General Brigham, burnout went from 50.6% to 29.4% at six weeks and held at 84 days. (JAMA Network Open, 2025) Both of those studies have ambient-vendor involvement in their authorship or institutional disclosures, which is worth knowing without being disqualifying. Kaiser Permanente, running the largest deployment in the country, logged 2.6 million uses in its first year of deployment and calculated 15,791 hours of documentation time saved for users relative to non-users. (NEJM Catalyst, 2025)

Then, in 2026, the largest and most careful study yet — 1,809 adopters against 6,772 controls across five health systems and three different products — reported the number that should anchor any honest conversation: 13.4 fewer minutes of total EHR time and 16.0 fewer minutes of documentation time per eight hours of scheduled care. That is a 3% and a 10% relative reduction. And time spent in the EHR outside work hours did not significantly differ between adopters and controls. (JAMA, 2026)

Why the numbers disagree

Two reasons, and both are actionable.

Use is shallow. In that multisite study only 32% of adopters used the scribe in more than half their visits — and that group got roughly twice the EHR-time reduction and three times the documentation reduction. In an emergency department study, the tool was used in 11.2% of eligible encounters and only 38% of attendings ever adopted it at all. (Annals of Emergency Medicine, 2026) Averaging committed users together with people who tried it twice produces a number that describes neither.

Perception outruns measurement. At UCSF, 86.5% of scribe users believed their documentation time had dropped — and there was no significant association between what they believed and what their logs showed. (AJMC, 2026) That is not a reason to dismiss the perception; feeling less besieged is a real outcome with real retention consequences. It is a reason to be suspicious of any vendor, including us, who quotes satisfaction survey results as if they were clock readings.

The honest summary

Ambient documentation reliably makes the work feel better and reliably shaves time off note-writing. It has not been shown to reclaim the evening. A meta-analysis of 23 studies found a real pooled effect on documentation burden alongside a blunt assessment of the evidence base: no randomised trials, small samples, high heterogeneity, and signs of publication bias. (BMC Med Inform Decis Mak, 2025)

What this means if you run a small practice

Three practical conclusions follow from the evidence rather than from enthusiasm.

Adoption depth is the whole ballgame. The difference between a scribe used in 20% of visits and one used in 80% is the difference between a rounding error and an hour a day. Which means the question to ask in a demo is not "how good is the note" but "what makes a clinician stop using this on a bad Tuesday." Usually the answer is friction: a separate app, a delay before the draft appears, an editing experience worse than typing.

Documentation time is not the only clock running. In the log data, note writing is 2.3 of the 5.8 hours. Chart review, orders and the inbox are the other 3.5. A tool that only writes notes is addressing 40% of the problem. This is why we treat orders, coding and the inbox as part of the same job rather than three separate products.

Measure your own logs, not the case study. Every health system in these studies had an analytics team pulling before-and-after event logs. A five-physician practice usually has a gut feeling. If you can get time-to-close-encounter and after-hours minutes out of your current system, capture them before you switch anything — otherwise you will be arguing about vibes in six months.

How we've built around it

ChartVoyant is voice-first rather than voice-attached: the note, the orders and the proposed codes come out of the same spoken encounter, and the clinician edits and approves in one place instead of reconciling a generated paragraph against structured fields it didn't fill. That design is aimed squarely at the adoption-depth problem — the tool has to be the path of least resistance on the worst day of the week, or the measured savings collapse to the average.

We would rather tell you that ambient documentation is worth roughly a quarter-hour a day at typical use and considerably more at heavy use, than promise you your evenings back and have you discover the JAMA result on your own.