1 The phrase that carries the weight
Nearly every deployment of clinical AI rests on one sentence: a clinician reviews the output before it reaches a patient. It is the answer given to regulators, to buyers, and to clinicians asking what happens when the model is wrong. It is also, as written, not a design. It names a person and a moment and specifies nothing about either.
The experimental literature is unkind to the unspecified version. In a randomized laboratory experiment, 120 final-year medical students prescribing with incorrect decision support made 86.6% more prescribing errors than with no decision support at all.1 Twenty-seven radiologists lost 36.8 to 59.9 accuracy points on mammograms where the AI-supplied BI-RADS category was wrong, least among the most experienced readers.2 Forty-four physicians, all of whom had received AI-literacy training, scored 14.0 percentage points lower on diagnostic reasoning when given deliberately flawed model output (95% CI −19.7 to −8.3).3 And in the one documentation pipeline measured end to end, review cut speech-recognition errors roughly twentyfold in volume while the clinically significant share of the survivors did not fall: 5.7%, then 8.9%, then 6.4%.4
A sibling paper appraises that evidence in detail. This one takes it as given and asks the engineering question: if the reviewing step is the control, what must the control be?
Two findings shape the answer. Over-reliance produces two distinct failures, omission, missing what the system did not flag, and commission, acting on what it wrongly flagged; a gate has to be built against both.5 And automation bias, across 40 studies drawn from nine databases, appeared in single tasks rather than multitasking ones, typically diagnosis, and where verification complexity was high.6 The reviewer’s problem is not divided attention. It is that checking is expensive. The opposite failure is equally documented: acceptance of decision-support alerts fell about 30% per additional reminder in the same encounter, 87.9% of repeats of an already-overridden alert were overridden too, and override rates showed no association with general workload.7 Asking for confirmation more often does not buy more confirmation.
2 What an approval gate is
An approval gate is a state transition that (a) cannot be performed by the system, (b) is performed by a named person who holds the authority to perform it, (c) is recorded as its own event, and (d) has a default of not proceeding. All four clauses are load-bearing. A mechanism missing any one is something else, usually a notification.
The system cannot perform it. The transition has to be unavailable to the software under every condition, including the ones it is most confident about. A system capable of the transition that merely declines to has implemented a policy, and policies live in configuration, which drifts. In ChartVoyant the drafted note is not a note: it is a draft that waits, nothing signs itself, and the draft stays a draft until the confirmation step. What holds the gate shut is the state the entry is in, not a rule about who presses what.
A named person with the authority performs it. Attribution to a role, a queue, a workstation or a service account is not attribution. The signing event in ChartVoyant’s audit trail carries the signing clinician and the timestamp, the minimum being an identity that resolves to one human being. Authority is a separate requirement, because that person must also be permitted to make this transition on this record. California’s AB 3030 shows the legal form of the same idea: its disclosure duties for generative AI used in patient communications about clinical information do not apply where the communication is “read and reviewed by a human licensed or certified health care provider.”8 The exemption turns on a licensed person, not on a review having nominally occurred.
It is recorded as its own event. A gate that leaves no trace distinct from the transition it authorizes cannot be examined afterwards, and a control nobody can examine is an assertion. The approval has to be a first-class record: who, when, against which version, with which outcome. ChartVoyant writes every action, including the AI’s, to a log that can only be added to, each entry locked to the one before it, so an altered history fails verification. Drafting and signing are separate entries, and the drafting entry links to the transcript it came from.
The default is not proceeding. The transition does not occur on a timer, a queue drain, a session expiry, a retry or a deploy. If nobody approves, nothing happens, and the non-event is visible as a pending item. Default-proceed gates are the most common failure in this class, because they are indistinguishable from working gates on every day that nothing goes wrong.
3 Where gates belong, and where they do not
The central trade is not whether to have gates but how many. A gate on every action reproduces the alert-fatigue pattern, and the mechanism is specific: repetition, not busyness, destroys attention.7 Across 17 studies, drug-safety alert override rates ran from 49% to 96%, with the reviewers explicit that overriding is frequently justified and the error-producing conditions systemic rather than individual.9 Excessive warnings lead clinicians to ignore the vital ones.10 The foundational clinical-decision-support design paper put speed first and told implementers to ask for additional information only when they truly need it.11 Too few gates produce the opposite failure: consequential actions leaving the practice with nobody answerable for them.
Four criteria decide the question. Reversibility: does undoing the action restore the prior state, or only append a correction? Blast radius: how many records, people and downstream systems does it touch? Egress: does it leave the practice, reaching a payer, a patient, another provider or a registry? External accountability: is someone answerable for it outside the practice, under a license, a signature, or a submission made in their name? A yes on either of the last two is close to dispositive; the first two decide the rest.
| Surface | Reversible | Blast radius | Leaves the practice | Answerable outside | What the criteria imply |
|---|---|---|---|---|---|
| Clinical note | No; corrected by addendum | One patient, every later reader | Yes, on release | The signing clinician, by license | Gate (the draft waits) |
| Diagnosis and procedure codes | Yes, until submitted | Claim and medical-necessity record | With the claim | The billing provider | Gate, or joined to the note gate |
| Claim to the clearinghouse | Only by void or corrected claim | Payer, patient balance | Yes | The billing provider | Gate |
| Referral and fax filing | Yes, but a misfile is a disclosure | Another patient’s chart | No, unless misfiled | The practice | Gate on chart assignment |
| Scheduling an appointment | Yes | One slot | Only as a notification | No one | Record, do not gate |
| Patient message on clinical information | No, once sent | One patient, acting on it | Yes | The reviewing clinician | Gate |
| Posting an insurer payment | Yes, by reversal | The ledger | No | No one at posting | Record, do not gate (posted automatically) |
For referral and fax intake the gate belongs on one decision inside the action: filing is reversible, but assigning a document to the wrong patient is a disclosure that refiling does not undo. Patient-facing messages carry two legal hooks, AB 3030’s exemption for provider-reviewed communications8 and the Texas requirement of a clear, conspicuous, plain-language disclosure that AI is being used in the provision of health care services.12 Both presume a record of which human being stood behind the message.
What the criteria are not
These four criteria are engineering judgment, not a measured result. No trial has compared gate placements in an EMR, and the evidence base is thinnest where the criteria say gates matter most: of 519 health-LLM studies, 44.5% evaluated on medical examination questions, 5% on real patient care data, and 0.2% each on billing codes and on prescriptions.13 Criteria derived by reasoning should be labeled as such, and revised when somebody measures.
4 Making verification cheap
This section matters most, and it follows from the finding that automation bias tracks verification complexity rather than divided attention.6 If checking one assertion is expensive, the gate produces confirmation without verification, the failure it was added to prevent. Four properties make checking cheap.
The source is one movement away from the assertion. For any statement in a draft, the material it came from has to be reachable without leaving the review surface and without a search. In ambient documentation that material is the transcript, and the link has to be recorded rather than reconstructed: ChartVoyant’s drafting event links to the transcript it drew on. One movement is the target, because the alternative is a reviewer who checks the first two claims and accepts the rest. In 97 outpatient encounters with 388 blinded paired specialist reviews, 31% of ambient notes contained hallucinations against 20% of physician-drafted notes, at near-parity overall quality ratings.14 That combination is exactly the condition under which unaided reading fails.
The diff is visible, and ordered by risk rather than size. A reviewer shown a finished document reads for plausibility. A reviewer shown what changed reads for correctness. The ordering has a measurable target: in 12,999 clinician-annotated sentences the overall hallucination rate was 1.47%, of which 44% were graded major, and the major ones clustered in the Plan section at 21%.15 A diff that gives a reworded history of present illness the same prominence as a changed plan has sorted attention by text volume instead of by consequence.
The gate presents what changed, not only the result. Confident, fluent, internally consistent output is the normal case, and models do not signal when they extend a false premise. In 300 physician-validated vignettes each seeded with one fabricated element, six models elaborated on the fabrication rather than rejecting it at rates from 50% to 82.7%, 66% overall; a mitigation prompt reduced that to 44%, and setting temperature to zero produced no significant improvement.16 A supported assertion and an unsupported one therefore have to look different on the screen, because they do not sound different on the page. The same holds upstream, where transcription systems have produced fabricated phrases with no counterpart in the audio, disproportionately after long non-vocal pauses, though the measurements come from aphasia and control speech corpora rather than medical dictation.17
The reviewer can see when a value is unavailable. A blank is ambiguous: it may mean normal, not measured, not retrieved, or not applicable. Federal certification rules agree. The Decision Support Interventions criterion at 45 C.F.R. § 170.315(b)(11) requires, at paragraph (v), that identified users be able to access, record and modify source attributes and to see when a source attribute value is unavailable,18 a requirement adopted in HTI-1.19 It binds health IT developers rather than practices, and it is a transparency obligation, not a validation one: no federal body evaluates model performance under it. Its future is unsettled, since a December 2025 proposed rule would remove the artificial intelligence model-card requirements entirely, with no final rule located as of September 2026.20 The design argument does not depend on the rule surviving.
5 What a gate must not do
No pre-checked consent. A checkbox that arrives checked has already made the decision and relabeled it as the reviewer’s. It fails clause (d) while appearing to satisfy the definition, and it manufactures a record of an approval nobody gave.
No single confident recommendation presented without its basis. The statutory test for decision support that stays outside device regulation has four elements, and the fourth requires software “enabling such health care professional to independently review the basis for such recommendations that such software presents so that it is not the intent that such health care professional rely primarily on any of such recommendations to make a clinical diagnosis or treatment decision regarding an individual patient.”21 The current guidance is FDA’s final version of 29 January 2026, which superseded the September 2022 guidance that most secondary commentary still describes.22 Law-firm analyses of the revision report that it accepts a single recommendation where only one is clinically appropriate; the objection here is not to singularity but to a recommendation whose basis is absent. The element is about the basis being independently reviewable, not about a click: adding a button saying a clinician must confirm does not by itself satisfy it. The radiologists in the mammography study were shown a BI-RADS category, a conclusion with no visible reasoning attached.2 A ranked set with its supporting and contradicting evidence exposed is reviewable. A bare label is not, however many times it is confirmed.
No gate whose only options are accept and accept-later. Rejection has to be reachable in one action, recorded as an outcome in its own right, and has to leave the draft available for correction rather than destroying it. A surface offering only approval and deferral is a queue, and queues get cleared.
No batching that lets one action approve many. One act of attention must not produce many approvals. The temptation comes from the same place as everything in section 3: repetition exhausts reviewers, so a hundred pending items invites a control that clears a hundred.7 A batch approval records what did not happen, which is worse than an honest record of an automated action.
6 Measuring whether a gate is working
A gate is a control, and an unmeasured control is a belief. Four families follow, all derivable from an append-only event record without additional monitoring of the clinician. First, time-to-decision distributions, read at the fast end rather than the median: a mode at the minimum possible interaction time is the signature of clicking, and it appears long before it appears in outcomes. Second, edit rates, the proportion of drafts altered before approval, segmented by surface, clinician and draft length. Third, the location of those edits, the operational version of the finding that dangerous errors concentrate in particular sections of a note.15 Fourth, rejection rates, which should be neither zero nor noise. A rejection rate of zero means the gate is not a decision; a rate indistinguishable from random means the drafts are not worth reviewing.
What these measures do not establish
A gate with a 100% approval rate is indistinguishable, from the outside, from no gate at all, and none of these measures shows that a reviewer understood what they approved. They measure the shape of attention, not its content, and the end-to-end documentation study is the analogue: review cut error volume roughly twentyfold without reducing the clinically significant share.4 Showing that a gate catches the errors that matter requires evaluation on real cases between offline validation and large comparative trials, the gap DECIDE-AI was built for with its 27 reporting items.23 This section specifies what to collect; no measurements of ChartVoyant’s own gate behavior have been published, and nothing here reports any.
7 The cost
Gates are friction, and the friction is not a rounding error. A systematic review and meta-analysis of human–LLM collaboration named the supervision burden a “vigilance tax”; across 10 peer-reviewed studies, composite diagnostic and management scores improved by 4.88 percentage points (95% CI 0.65 to 9.12) with a 95% prediction interval crossing null, while factual error rates remained around 26% to 36%.24 In a randomized trial reported as a preprint, physicians with model access took 119.3 seconds longer per case (95% CI 17.4 to 221.2).25 In a randomized trial of diagnostic reasoning, physicians using an LLM scored no better than physicians without one, an adjusted difference of 2 points (95% CI −4 to 8; P = .60), while the model alone scored 16 points above the control arm.26 A later trial found that where the model sat in the sequence of work changed the result.27 Review time is real work, and buying it back by weakening the review is the failure this specification exists to prevent.
The honest statement is this: a system with proper approval gates is slower than the same system without them, and that is the intended trade rather than a defect to be optimized away. The tension with the design literature’s first commandment, that speed matters most,11 is why the placement criteria exist. The response to friction is fewer gates in better places, with cheaper verification at each. It is never a gate that proceeds on its own.
8 The properties, as a test
The specification reduces to ten questions, answerable about any system that claims a human is in the loop.
- Can the system perform the transition itself under any condition, including a timeout, a retry or a deploy? If so, there is no gate.
- Does the record name one person rather than a role, a queue or a service account?
- Did that person hold the authority for this transition on this record, and is that checked rather than assumed?
- Is the approval a separate event from the transition it authorizes, on a record that cannot be rewritten?
- What happens if nobody acts? If the action proceeds, the default is wrong.
- How many seconds and how many movements does verifying one assertion against its source take?
- Does the surface show what changed since the last approved version, ordered by consequence, or only the current result?
- Can the reviewer distinguish a missing value from a normal one?
- Is the basis of a recommendation reviewable independently of the recommendation?
- Is rejection a first-class outcome with its own record, and what is the observed rejection rate?
None of these asks whether the model is good. A better model does not remove the need for a gate: the failures above all occurred with output good enough to be believed. A gate is a statement about who is answerable when the output is wrong, plus a specification for keeping that person’s attention affordable. The difference between a real gate and a ceremonial one is visible from outside, in the ten answers.