Articles
A research library on AI and the health record.
Twenty-one papers, written to be cited rather than skimmed. Some set out how an AI-native EMR is actually built; others review the published literature, weighted by study design, and correct a number of figures this field repeats without checking. A third group reads the rules an independent pain practice works under — PDMP checks, EPCS, drug testing and Medicare coverage — against the evidence behind them. Every claim carries a reference, and where the evidence will not support a claim the papers say so.
This is not the Insights section. Insights is shorter and written to be read once. Articles is a reference library, for a clinician, engineer, administrator or researcher who has to form a defensible view of what artificial intelligence does inside an electronic health record, and of the rules a practice's record has to satisfy. The papers assume a reader who will check the citations. Where a source published no abstract, it is cited for its existence and scope and no number is taken from it; where a widely repeated figure turns out to rest on a press release, a pleading or a vendor extrapolation, the papers say which. None of it is legal, billing or clinical advice.
The papers
21 papers · last reviewed September 2026-
01
A write-path architecture for an AI-native EMR
The decisive design choice is not the model but the write path: what may
enter the record, in what state, on whose authority, and with what provenance. The
architecture makes generated text attributable and reviewable; it does not make it
accurate.
ArchitectureOriginal technical description
28 references · 14 min -
02
Approval gates: a specification for human decision authority in EMR automation
“A clinician reviews it” is not a design. A gate is defined as a
state transition the system cannot perform, made by a named person with authority,
recorded as its own event, and defaulting to not proceeding — with criteria for
where gates belong and how to stop them decaying into clicking.
Oversight designOriginal design specification
27 references · 14 min -
03
Tamper-evident clinical records: a hash-chain specification for AI-authored documentation
An append-only event log in which each entry is bound to its predecessor,
why an EMR that lets a machine write needs one, and a candid account of what
tamper-evidence does not prove — including where the blockchain literature gets
the trade-off wrong.
Record integrityOriginal technical specification
23 references · 14 min -
04
Retrieval over the chart: assembling clinical context from standards that were not built for it
Before a model drafts anything, software decides what it may read. The
interoperability standards it reads through were designed for document exchange, not
question answering — including a three-version gap between the published USCDI and
the version actually adopted in regulation.
Data and retrievalArchitecture and design rationale
28 references · 14 min -
05
Ambient documentation and the clinician’s day: a review of the evidence
The randomized evidence shows a real effect that is small and specific to
the product: one tool cut time in the note by 9.5% while a second tool, randomized in
the same trial, changed nothing. The familiar figure of about an hour saved per day is
supported by none of the studies reviewed.
Documentation burdenSystematic appraisal of the literature
28 references · 15 min -
06
Evaluating clinical language models: benchmarks, blind spots, and what the evidence supports
Of 519 studies of language models in health care, 5% used real patient care
data and 44.5% used examination questions. This corrects three claims the underlying
papers do not support, and sets out what an accuracy figure needs before it carries
information.
EvaluationCritical review of evaluation practice
28 references · 15 min -
07
Human oversight of clinical AI: what the automation-bias literature actually shows
Review reliably reduces the volume of errors and does not reliably reduce
the share that are clinically significant. Where the machine is wrong, reviewers
perform worse than unaided clinicians, and training and experience attenuate the
effect without removing it.
SafetyNarrative review with design implications
27 references · 15 min -
08
Automation in the revenue cycle: coding, denials, and prior authorization
What is measurable about coding accuracy, denial rates and prior
authorization burden — and why so many of this field’s most quoted numbers
cannot be traced to a source, including several that circulate through three trade
associations and read as three studies.
Revenue cycleEvidence and regulatory review
28 references · 15 min -
09
Regulating AI in certified EHR technology: what actually binds a vendor
Health IT certification, FDA device regulation, and a voluntary assurance
layer that is not regulation at all, with exact citations and dates. The certification
criterion compels disclosure and documented risk management. It sets no performance
threshold, and neither does anything else that reaches an EMR.
RegulationRegulatory review
28 references · 15 min -
10
After go-live: dataset shift, silent failure and the monitoring of clinical AI
Deployed models degrade slowly through calibration drift and abruptly on dated events: the ICD-10 transition shifted about 1 in 6 diagnostic categories by 20% or more in two of three classification systems. Stable discrimination can mask failing calibration, and the certification rule requires monitoring to be described, not to find anything.
Model monitoringNarrative review with operational implications
29 references · 14 min -
11
Prompt injection and the tool-using clinical assistant: a threat model
Models tested on medical tasks followed injected instructions at high rates, in one study in 94.4% of injected simulated patient dialogues. Filters and detectors lowered attack rates on fixed tests and failed against adaptive attackers. The controls that hold limit what a compromised assistant can do, not what it can be told.
SecurityThreat model and literature review
29 references · 15 min -
12
Bias in clinical algorithms, and the federal rule that reaches the practices using them
Many of the most cited cases of algorithmic bias arose from labels, measurement and training data, not race variables. The Section 1557 rule’s identification duty is triggered by input variables, which catches a pain clinic’s opioid risk tool but not a cost-trained risk score. The rule remains in force as of September 2026.
EquityEvidence and regulatory review
28 references · 15 min -
13
Telling patients a machine wrote it: the evidence on disclosure, and the laws that now require it
Patients say they want to be told, an AI disclosure costs 0.09 to 0.13 points of satisfaction on a five-point scale, and fuller explanation cut ambient-recording consent from 81.6% to 55.3%. California exempts clinician-reviewed drafts from disclosure; Texas does not; no Tennessee provider disclosure statute was found.
TransparencyEvidence and regulatory review
26 references · 14 min -
14
AI-drafted replies to patient messages: what the deployments measured
In the one randomized study located, drafts raised physicians’ read time by 21.8% without changing reply time. Burden relief is self-reported and strongest among nurses, the preference studies rewarded length, and physicians missed two in three known errors in simulated drafts. None of the deployments reviewed audited what was sent.
Patient communicationSystematic appraisal of the literature
24 references · 14 min -
15
PHI and the model: what HIPAA requires when a language model touches the chart
A model vendor that receives PHI is a business associate, not a conduit, and HIPAA prescribes what its agreement must contain. It sets no retention period during the contract and is silent on training, hosting location and ownership of output, and HHS certifies nothing as HIPAA compliant.
PrivacyRegulatory review
29 references · 15 min -
16
Ransomware and the independent practice: what the breach record shows
In 2024 business associates reported 16% of large breaches but 85% of affected individuals, mostly from one clearinghouse attack. Attacks measurably disrupt hospital care; harm to practices is documented here only by surveys. Recommended controls rest on guidance and case evidence, and the one trial located, of training, found little effect.
SecurityEvidence review
29 references · 14 min -
17
Prescription drug monitoring programs: what mandated checking has and has not been shown to do
Must-access mandates reliably reduce opioid prescribing and multiple-provider episodes, partly among patients they were not aimed at, while compliance at the prescription stays low where measured. Overdose effects are mixed, and several studies find more heroin deaths. Proprietary risk scores lack clinical evaluation. Tennessee’s exact check schedule is stated and a superseded version corrected.
Controlled substancesEvidence appraisal with regulatory analysis
26 references · 15 min -
18
Electronic prescribing of controlled substances: what the DEA rule requires, and what it proves
An EPCS signature proves that an identity-proofed registrant signed specific contents that could not then change; it does not prove the prescription was appropriate. The DEA rule is still interim, several standards it names are withdrawn, and the national study found wider EPCS use was not associated with less opioid prescribing.
Controlled substancesRegulatory review
30 references · 14 min -
19
The 2022 CDC opioid guideline in the record: from dose thresholds to decision support
Voluntary thresholds from the 2016 CDC opioid guideline became statute, pharmacy edits and EHR alerts; the 2022 guideline moved them out of its recommendations. EHR defaults reliably change quantities prescribed. Concordance alerts are seldom tested against patient outcomes, and average daily MME varies threefold with its definition.
Clinical guidanceGuideline and evidence review
27 references · 15 min -
20
Urine drug testing in chronic pain care: what the evidence, the guidelines and the coverage rules support
A systematic review found no study of whether risk-mitigation strategies such as urine drug testing reduce overdose or addiction, and the 2022 CDC recommendation carries its lowest evidence grade. Standard screens miss oxycodone and fentanyl. Medicare in Tennessee pays for individually justified, risk-tiered testing, and the LCD’s example risk tool had 18 patients in its original low-risk group.
Drug testingEvidence, guideline and coverage review
26 references · 14 min -
21
Interventional spine procedures and the Medicare coverage determinations: what the record has to show
Palmetto GBA’s facet and epidural LCDs read as a documentation specification: named scales at baseline, the same scale after every block, 80% relief on two blocks before ablation, rolling session counts. Federal audits find the gaps in records and claims; in 2006, 8% of facet services had a medical-necessity error.
CoverageRegulatory and evidence review
26 references · 15 min
How to use this library
The papers are written to stand alone, but they were built as a set and they argue with each other in places. Search reaches every paragraph and every reference.