The expensive version of billing is reconstruction

A typical independent practice does not have a billing problem in the sense of “we forgot to send the claim.” It has a reconstruction problem. The visit happens in one system. Someone later re-reads the note, picks codes, keys them into another system, hopes the diagnosis on the claim matches the diagnosis in the chart, and finds out weeks later whether it did.

That gap is where money leaves. Practice leaders asked to name the single biggest leak in their revenue cycle put denials and appeals first, at 48%, with front-end processes second at 23% and billing and collections at 14%. (MGMA Stat, 2026) The ranking is slightly misleading. Denials and front-end processes are the same failure observed at two different moments — a fact we wrote about at more length here.

What that piece did not spend time on is the work that happens after a clean visit: assembling the 837, scrubbing it, submitting it, posting the 835, and working whatever comes back. That work is large, mostly mechanical, and still done by people sitting in a second application.

A biller reconstructing a claim from a narrative is doing the same job twice. AI that helps them do it faster is still doing it twice.

The numbers on AI-in-billing are not encouraging, and that is informative

An HFMA–FinThrive poll of 101 organisations found that 63% already use AI or automation somewhere in the revenue cycle, and that only 15% of those report a positive return. The leading obstacle was not model quality. It was infrastructure: 51% cited IT limitations, 43% integration with existing systems, 42% difficulty demonstrating ROI. (HFMA, 2025)

Read that as a product finding, not a model finding. If the claim has to be exported, mapped, and re-imported before a model can see it, you have bought a translation layer. Translation layers have integration costs that eat the time they were supposed to save. CAQH’s 2025 Index is consistent with this: the industry avoided an estimated $258 billion through electronic transactions, and still has about $21 billion sitting in the remaining manual and portal work — eligibility, claim status, prior authorization, payment. More than half of health plans now use AI in administrative workflows; about a quarter of provider organisations do. (CAQH Index, 2026)

The per-transaction arithmetic has been stable for years, which is itself a tell. A manual eligibility check costs the provider about $7.97 against $2.18 fully electronic. A prior authorization is $10.97 against $5.79, and 22 minutes against 11. A claim-status inquiry by phone is the longest transaction CAQH measures. (CAQH Index, 2023) Those gaps do not close because someone types faster. They close when the transaction is never re-keyed.

Where a language model is actually good in the revenue cycle

The published evidence on models choosing medical codes is poor, and we have already written that piece: every leading model scored below 50% at reproducing ICD and CPT from clinical descriptions, and expert human coders disagree with each other on 56% of office encounters. Autonomous coding is the wrong job.

The jobs language models are good at sit one layer down, and they compound:

1. Organizing the medical decision making so a person can pick the level

Problems addressed, data reviewed, risk of management — pulled from what was said and done, laid out against the two-of-three rule, each element traceable to the encounter. That is summarisation. Summarisation is the thing the models do well. Classification of an E/M level is the thing they do badly. The correct split is: the model shows the evidence; the clinician attests the code.

2. Assembling the claim from the attested visit, not from a reread

Once the diagnosis and procedure codes are attested, an 837P is a structured document with known loops: subscriber, billing provider, claim, diagnosis HI, service lines. Generating that from the encounter record is not a language-model problem. It is a mapping problem. The AI contribution is making sure the mapping has something true to map from — a signed note, attested ICD-10-CM, attested CPT/HCPCS, coverage that was checked today, an NPI that is real.

3. Scrubbing before the clearinghouse does

Missing diagnosis pointers, a CLM02 total that does not equal the service lines, a placeholder EIN on a live wire, a procedure with no supporting ICD. These are deterministic edits. A model is useful for explaining them in English and suggesting the chart fact that would close them. It should not be useful for inventing the fact.

4. Reading what comes back

An 835 is a payment file, not a letter. Parsing it and posting it to the right claim is mechanical, and doing it by hand is how patient balances go stale. A denial letter, by contrast, is a letter. Extracting the reason code, the appeal deadline, and the one sentence that actually matters is a language-model job. Drafting an appeal that cites the note already in the chart is the same job one step further. With 57% of Medicare Advantage denials ultimately overturned and roughly 40% never resubmitted, the abandonment rate is the number that costs a practice money — and abandonment is a time problem. (Health Affairs, 2025)

The model should never decide what to bill. It should make it expensive, in time, to send a claim that does not match the visit, and cheap, in time, to work one that comes back.

What this looks like when billing lives in the EMR

ChartVoyant’s billing system is not a clearinghouse bolted on after sign-off. It is the same record the visit was documented in.

Insurance is captured as structured fields at check-in, with the card image kept as evidence. Eligibility runs against the plan being presented today. The note is written from the visit; the diagnosis and procedure codes are proposed from the medical decision making and attested by the clinician before anything downstream can use them. On that attestation, the system captures charges, assembles an 837P from real ICD-10-CM and HCPCS, runs a scrub, and parks the claim on a worklist. Remittances post from a parsed 835. Denials land with an owner, not in a shared inbox.

The AI sits in the joints of that pipeline, not on top of it:

  • During the visit it organizes the MDM and shows the E/M level that follows, with the one next to it, so the clinician is choosing rather than accepting.
  • At sign-off it has already linked the visit diagnosis to the claim diagnosis. A whole class of medical-necessity denials is just that linkage being done twice, differently.
  • After a denial it drafts the appeal from the note, the order and the reason code, so working the queue is review rather than authorship. Forty minutes of reconstruction becomes a few minutes of reading.

None of those steps require the model to be a coder. They require the model to be looking at the same chart the biller is looking at, which is only possible if there is one chart.

What we will not claim

We will not quote a first-pass clean-claim rate we have not measured on real claims, and we will not advertise autonomous coding accuracy. The published benchmarks on the latter point the other way, and CERT data on office visits says the dominant error is the level, not missing documentation — which is exactly why a person still has to pick it. Production eligibility and 837P claims are configured to go through Stedi under a signed BAA. The public demo stays on synthetic claims. We still have not measured a clean-claim rate on real claims.

The test that actually matters

Ask a vendor to show you, on one visit, the moment the claim comes into existence. If that moment is a person opening a second application and reading the note, you are looking at reconstruction with a nicer interface. If that moment is sign-off — codes attested, 837 already shaped, scrub already run, denial path already owned — you are looking at a billing system. The AI is what makes the second version cheap enough to run on every encounter rather than on the ones someone had time for.