1 The write path, not the model
Most products described as AI-native electronic medical records are a model attached to a record system designed on the assumption that a human typed everything. Its schema, permissions, audit log and notion of authorship all predate the new author. The generative component is given somewhere to put its output, usually a text field that already existed, and engineering attention goes to prompts and evaluation.
The research literature has the same emphasis. A systematic review of 519 studies of large language models in health care found that only 5% used real patient care data while 44.5% evaluated on medical examination questions, and that the administrative work record systems actually mediate was close to absent, at 0.2% of studies each for billing codes and prescriptions.1 That is a literature about model capability, not records.
This paper takes the opposite order. In an EMR the record is the product: a legal instrument, the substrate of a claim, evidence in a liability case, and what the next clinician relies on. The first question is not what a model can generate. It is where a model may write, what state its output occupies on arrival, who may move it out of that state, and what has to be true of the record before anything is committed. That set of rules is the write path.
This paper describes one write path, as implemented in ChartVoyant, with the reasoning behind each constraint. The evidence comes from the published literature, not from measurements of this system: no performance claim about ChartVoyant appears here, because none has been published.
2 Author, state, and authority
Every entry carries three properties. They are independent, and most failure modes in AI documentation collapse two of them into one.
Author
The author is the named agent that produced the content, person or machine, recorded on the entry rather than inferred from the field the text landed in. This is the Agent of the W3C provenance model, distinct from the Activity that produced the Entity — described here from PROV-Overview, a Working Group Note; the normative document is PROV-DM.2 FHIR splits the work: Provenance covers generation, AuditEvent covers usage, so recording authorship in one and expecting it to answer access questions leaves a gap. Provenance is Trial Use, not normative.3
State
State is the entry’s standing in the legal record. Two are sufficient: draft, recorded and durable but not part of the record a clinician is accountable for, and attested, meaning a clinician has taken responsibility for it. State belongs to the entry, not to whoever wrote it: human-typed text is unattested until signed, exactly as machine-drafted text is.
Authority
Authority is the right to move an entry between states. It belongs to a person, scoped to a role and a practice, never to an author by virtue of authorship. FDA’s electronic records rule, which governs FDA-regulated records rather than ambulatory EMRs, states most clearly what a signature must carry: the signer’s printed name, the date and time, and the meaning of the signature, whether review, approval, responsibility or authorship.4 It is an unflattering comparator.
The applicable American rule is thinner. Medicare’s signature guidance states that a scribe documenting entries on a provider’s behalf need not sign or date the documentation; the treating clinician’s signature affirms that the note adequately documents the care provided.5 Applied to a machine drafter, that puts the whole attestation burden on the signer and requires no disclosure of what produced the text.
Why collapsing them fails
Collapse author into authority and the component that writes becomes the component that decides. Collapse state into author and the record acquires a false rule, that machine output is provisional and human output final, which breaks the moment a clinician edits a draft. Collapse author into state and the record forgets who wrote a sentence at the moment it is signed, the most damaging of the three: attestation is exactly when provenance becomes worth having.
None of the three is supplied by regulation. HIPAA’s audit controls standard is one sentence requiring mechanisms that record and examine activity in systems holding electronic protected health information, with no content list, no retention period and no review cadence, and the mechanism for corroborating that information has not been improperly altered is Addressable, the weakest tier.6 Reviewing activity records is Required, but “regularly” is undefined.7 What specificity exists comes from certification8 and its audit-log content standard.9 A proposed Security Rule overhaul was still a proposal as of September 2026.10
3 Draft as a first-class state
It follows that the model’s output is a durable, recorded object that is not part of the legal record until a clinician acts on it. In ChartVoyant the transcript fills during the visit and a draft is produced alongside it. Nothing signs itself: the draft stays a draft until a clinician reviews, signs and locks it into the record.
Three event types carry the visit across that boundary. visit.transcript.appended records speech captured during the visit. ai.note.drafted records the machine-written draft and links to the transcript it came from. note.signed records the signing clinician and the timestamp. Every action, including the model’s, goes to a log that can only be added to, never edited or deleted, and each entry carries its own digest and a pointer to the previous entry’s digest, so altering history is detectable. Records, files and audit logs are separated per practice; applied here, that yields a per-practice chain.
The construction is old and needs no distributed ledger: Haber and Stornetta’s linking scheme chains documents through cryptographic hashes, making back-dating computationally infeasible.11 Chaining seals entries against one another, but the times in the chain are the system’s own; an external time reference is a separate guarantee, obtained from a timestamping authority that signs over a hash without seeing the data behind it.12
Append-only is not immutable, and the difference matters clinically: records must stay correctable, since clinicians retract, addenda are added and patients may request amendment. What is needed is append-only with tamper-evidence.
Why the draft is written to the log
Holding a draft in memory until a clinician accepts or discards it is simpler, cheaper and wrong. An unreviewed suggestion that disappears cannot be audited: if the only drafts that survive are the ones a clinician approved, the log records the system’s successes and nothing else. A rejected draft is evidence too.
The reviewer’s decision also means nothing if the object decided upon was not preserved. An attestation is a statement about a specific artifact, and a signature over a draft that no longer exists, or that the reviewer mutated in place, is a signature over nothing recoverable. Keeping the draft as one entry and the signature as another is what allows the difference between them to be reconstructed, which is the only way to establish what the clinician changed.
4 Why the transcript is the anchor
The link from ai.note.drafted back to visit.transcript.appended is the load-bearing part of the design. Review is only a control if verification is cheap.
Lyell and Coiera analyzed 40 studies from nine databases covering 1983 to 2015, only 6 of them healthcare-focused. Automation bias was not principally a multitasking artifact: it appeared in single tasks, typically diagnosis rather than monitoring, and where verification complexity was high.13 That names a variable the designer controls.
Leaving verification expensive has a measured consequence. In a randomized experiment with 120 final-year medical students, incorrect decision support raised total prescribing errors by 86.6% relative to no support at all, while correct support reduced them by 58.8%.14 In a single-blind randomized trial of 44 physicians who had all received AI-literacy training, deliberately flawed language-model output cut diagnostic reasoning scores by an adjusted −14.0 percentage points (95% CI −19.7 to −8.3; P<.0001).15 Training did not protect them. Putting the source document one click from the generated sentence acts on the variable the evidence identifies.
The failures worth designing against are counted. Across 450 clinical note and transcript pairs, clinicians annotated 12,999 sentences for hallucination and 49,590 for omission: 1.47% of sentences were hallucinated, 44% of those graded major, and 3.45% were omissions. Major hallucinations clustered in the Plan section, at 21%.16 The highest-risk fabrications sit in the part of the note that says what happens next.
| Study | Unit of measurement | Reported rate |
|---|---|---|
| Asgari et al. 202516 | Sentences of a generated note | 1.47% hallucinated (44% major); 3.45% omissions |
| Palm et al. 202517 | Notes containing at least one hallucination | 31% ambient vs 20% physician-drafted |
| Omar et al. 202518 | Vignettes seeded with one fabricated element | 66% overall; 50–82.7% by model |
| Koenecke et al. 202419 | Transcriptions of non-medical speech corpora | ~1% held a fabricated phrase; 38% of those harmful |
| Zhou et al. 201820 | Words of dictated clinical documents | 7.4 errors per 100 words at recognition output |
Grounding does not make a draft correct. In an adversarial study of 300 physician-validated vignettes, each seeded with one fabricated element, six models elaborated on the fabrication rather than rejecting it, 66% of the time overall; a mitigation prompt cut that to 44%, and temperature zero produced no significant improvement.18 A model handed a false premise builds on it.
The anchor is itself a recording rather than ground truth. An audit of a widely used speech-to-text system found entirely hallucinated phrases in roughly 1% of transcriptions, 38% of them containing explicit harms, disproportionately triggered by longer non-vocal pauses.19 Those corpora were aphasia and control speech, not medical dictation, so that is not a medical-setting error rate. The design rests on something narrower: the transcript is the recorded input the draft came from, and what a reviewer checks a sentence against.
5 What the write path must refuse
The specification is easier to state as prohibitions; four refusals define it.
1. No silent writes. Every entry names its author, and a machine author is named as such rather than inheriting the identity of the session or the clinician it ran for. Certification requires actions on electronic health information to be recorded against an adopted content standard,9 but nothing requires a machine-drafted narrative to be identified as machine-drafted: the source-attribute disclosures reach decision support interventions and stop there.8 A system closes that gap by making the drafting event a first-class event type with its own author, which is what ai.note.drafted is.
2. No model-initiated state transitions. Only a licensed clinician may move an entry from draft to attested, including entries the model authored. Clinical suggestions are drafts; the model decides nothing on its own. A wrong suggestion a reviewer accepts is worse than no suggestion,14,15 so no path to the record may bypass a reviewer. Because the signing clinician’s attestation is what makes a note authoritative,5 any mechanism producing an attested entry without a signer manufactures authority no one granted.
3. No post-hoc edits that leave no trace. Corrections are new entries and nothing is overwritten in place. The record already has a provenance problem predating generative systems: across a decade at one academic medical center, median outpatient note length rose 60.1% and notes written in 2018 averaged 29.4% directly typed text.21 Text that arrived by copy or template carries no record of its origin, and adding a generative author to a record that cannot say where its sentences came from turns a quality problem into a categorical one.
4. No reconstruction of an encounter from a downstream artifact. The claim, the referral letter and the printed chart are outputs of the record; none is the encounter. Printed renderings differ in formatting and chronology from the record beneath them, which is why recreating an accurate timeline in a liability case is specialist work.22 Permitting downstream reconstruction guarantees two versions of the encounter and no way to say which is authoritative.
| Event | Author | State on write | Who may advance it |
|---|---|---|---|
visit.transcript.appended | Capture of the visit | Recorded, not asserted | No one; appended to, never rewritten |
ai.note.drafted | The model, named, linked to the transcript | Draft | The treating clinician, by signing |
note.signed | The signing clinician, with timestamp | Attested | No one; a correction is a new entry |
claim.submitted | Billing, projected from the attested note | Derived | The payer; its response is a new entry |
payment.posted | Posting, from the payer’s remittance | Derived | No one; adjustments are new entries |
6 Downstream artifacts are projections
The last two rows of Table 2 follow from the first three. claim.submitted records an 837P sent to a clearinghouse and payment.posted records an insurer payment against it, both in the same append-only chain as the events they descend from. A claim, an order or a letter is a projection of an attested encounter, not a reconstruction of it.
Two consequences follow. The design requires that a claim be projectable only from an attested encounter, because otherwise the projection has no source. And the code set that reaches a claim is the set a clinician confirmed: ChartVoyant’s drafted output includes suggested codes, illustrated publicly with ICD-10-CM M54.16 and G89.29, and those are part of the draft, subject to the same confirmation as the narrative. Billing codes were the subject of 0.2% of studies in the review cited above.1 That is a reason to be conservative here, not to treat coding as easy.
7 What this architecture costs
The costs are real, and several are permanent rather than implementation debt.
Storage and event volume. The design requires every draft to be retained, including rejected ones, alongside the transcript it came from. The log grows with activity, not with the size of the chart, and it never shrinks.
Latency between speech and a usable note. The draft appears during the visit; the note does not exist until someone signs it. No model improvement removes that interval, because the interval is the control.
A review burden that does not go away. A randomized study of AI-generated draft replies to patient messages found read time rose 21.8% (95% CI 5.2% to 41.0%; P=.008) with no significant change in reply time (−5.9%; P=.33).23 A meta-analysis of human and language-model collaboration found a composite improvement of 4.88 percentage points (95% CI 0.65 to 9.12) with a prediction interval crossing null, factual error rates persisting at 26% to 36%, and collaboration not universally outperforming the model alone; the authors call the supervision cost a “vigilance tax.”24 That tax is what this write path charges.
A benefit that is not categorical. A pragmatic randomized trial of 238 outpatient physicians across 14 specialties found one ambient product reduced time-in-note by 9.5% (95% CI −17.2% to −1.8%; P=0.02) while a second showed no significant change (−1.7%; P=0.66).25 Architecture does not determine whether a product saves time. Review time is not free either: each additional hour a primary care physician spends on documentation is associated with a 7.1% decrease in the likelihood of accessing outside patient records.26
A constraint on what can be automated. Auto-signing a normal note and auto-submitting a clean claim are unavailable by construction. That is the intent, and still a cost. The gap the review step exists to catch is measurable: across 5 standardized primary care cases rated by 30 blinded raters, notes from 11 AI scribe tools scored lower than notes from 18 human note takers on every case and on all 10 modified PDQI-9 domains.27 Those cases were simulated, so the gap is an upper bound.
Limitations of this design
The write path buys attribution and reviewability at a price paid in storage, event volume, latency and clinician attention, and it forecloses automations a less constrained system could ship. None of those costs is recovered by a better model, because none is caused by the model.
8 What it does not solve
This architecture does not make the model accurate. It makes the model’s output attributable and reviewable. Those are different properties, and the second does not imply the first.
The sharpest evidence for the distinction predates language models. Across 217 dictated clinical documents, the error rate per 100 words fell from 7.4% in raw speech-recognition output to 0.4% after transcriptionist editing and 0.3% in physician-signed notes. The proportion of remaining errors that were clinically significant did not fall: 5.7%, then 8.9%, then 6.4%.20 Review removed errors; it did not preferentially remove the dangerous ones. A write path routing generated text past a reviewer inherits that result exactly.
The log has limits of its own. Across 85 studies using EHR audit logs to study clinical activity, only 19 (22%) validated their results and 9 (11%) validated against direct observation.28 An audit chain is strong evidence of what the system recorded, not of what happened in the room. It makes a record’s history checkable, not its contents true.
No measurement is reported here
This paper reports no accuracy, time-saving, documentation-burden, coding or denial-rate measurement for ChartVoyant, because none has been published. What is described above is a specification, not a demonstrated outcome. ChartVoyant’s own security documentation states that no system is perfectly secure and that the company does not claim otherwise; the same register applies here.
9 Conclusion
The model in any AI-native record system will be replaced, probably within a year. The record will not. That asymmetry is the argument for specifying the write path first and treating the model as one more author subject to it: named, producing output in a state it cannot leave on its own, anchored to the input it came from, in a log that can be appended to and not rewritten.
What such a design yields is a set of answerable questions. What did the model produce. What was it given. Who reviewed it, and what did they change. What was sent downstream, and from which attested entry. A system that cannot answer those does not solve them by choosing a better model; one that can has not thereby made the model right.