Case studyApplied AIAll work

A model that asks instead of guessing

An AI coding engine for a US healthcare revenue-cycle company. It reads anaesthesia charts, assigns every billing code with its reasoning, flags what the record leaves out — and puts each code in front of a coder to accept or reject.

Client A US healthcare revenue-cycle companyEngagement Three engineers and a business analyst · 2025 — presentPractice Applied AI

Before a hospital or a practice can be paid for a procedure, somebody has to turn the clinical record into codes: what was done, why, under what kind of anaesthesia, for a patient in what condition. Payers pay against the codes, not the notes. The client does that coding for healthcare providers, and an experienced coder takes ten to fifteen minutes a chart.

That is the constraint the business runs on. Coders are skilled, scarce and expensive; volume arrives in batches; and a code that is wrong costs money in one direction or a compliance problem in the other.

We built an AI coding engine, starting with anaesthesia charts. It is live and integrated with the client's revenue-cycle systems. It reads the chart, assigns the codes, explains its reasoning and the alternatives it considered, raises alerts where the documentation is missing something the code depends on — and hands every code to a coder, who accepts or rejects it.

Charts coded

19,000+

Anaesthesia charts processed by the engine so far.
Per chart

About 75 seconds

Against ten to fifteen minutes by hand. The coder's job becomes checking the engine's argument rather than building it.
Per day

5,000–7,000

Charts a day in bulk processing.

What makes this hard

  • The input is a scanned record, not a form. A single case runs to a dozen pages — consent, pre-operative assessment, the anaesthesia record itself — and parts of it are handwritten.
  • The answer is several codes that depend on each other. The procedure code and up to eight ancillary codes, medical direction, the anaesthesia code, a physical-status modifier, the mode of anaesthesia, any qualifying circumstances, and up to ten diagnosis codes behind them.
  • The rules move. The code sets are revised every year, and individual payers add rules of their own.
  • The data is protected health information. Patient records cannot be sent anywhere casually, and every design choice starts from that.
  • A confident wrong answer is the worst answer. A coder can work with "unclear". A plausible code that happens to be wrong gets billed.

The engine's most important output is not a code. It is the reason for the code, and the admission when the record does not contain one.

From chart to codes

One chart through the engine
1De-identifyPatient names, record numbers and account numbers are removed before any model sees the chart.
2ReadCharacter recognition turns handwriting and scanned pages into text.
3ExtractA smaller model pulls out what is relevant to coding, so the reasoning step works from the clinical facts rather than twelve pages of forms.
4RetrieveThe crosswalks between procedure and anaesthesia codes, the modifier and qualifying-circumstance logic, and the client's own rules — which codes combine, how diagnoses are sequenced — are indexed in a vector store, and what matches this chart is retrieved alongside it.
5ReasonA reasoning model assigns the codes from the extracted record, the retrieved rules and the coders' own methodology — and records why, and what else it considered.
6ValidateA coder sees each code with its confidence and reasoning, and accepts or rejects it. Nothing is final until a person says so.
7LearnCorrections go back in as training, so the engine learns this client's practice rather than coding's average.

A reasoning model rather than a plain language model was a measured choice: on documents like these it gave more accurate answers. The cost of that choice is that it needs clean text to reason over, which is why reading and extraction are separate steps in front of it rather than one model doing everything.

Which reasoning model was measured too. The engine runs on o4-mini, chosen after comparing it with larger models — GPT-5.1, o4 and Claude models among them. At thousands of charts a day, the smallest model that is accurate enough is the right one, and the comparison is what tells you which that is.

Showing the work

Every code comes back with two things a coder can check: the reasoning behind it, and the other codes the model considered and why it rejected them. That turns review from re-coding the chart into checking an argument — faster, and far easier to audit.

It also exposes where the model is weakest, which is the point.

When the record does not say

The instructive failure came early. For some procedures, the correct code depends on a detail — the material used, say — that the notes simply do not state. Asked to code such a chart, the model narrowed the answer to two codes and explained that the record did not settle it. Forced to choose one, it picked the more common option.

That is exactly what a good coder would not do. A good coder sends a query back to the clinician. The typical answer is not the answer for this patient, and billing it is a guess dressed as a decision.

So the engine is built to say so. Where the documentation is missing something a code depends on, it raises a missing-data alert rather than filling the gap with the most likely value — and the coder, not the model, decides what happens next.

Measured before it was trusted

The proof of concept was designed to prove the engine on the client's own history before it touched live work, in five rounds.

  1. Hold out a test set.A share of charts with verified coding is set aside from the start and never used for training, so accuracy is always measured on charts the model has not seen.
  2. Measure the untrained model first.The first round codes without any of the client's charts, so every later improvement is measured against what the model knew on its own — not against nothing.
  3. Train, then feed back.Later rounds add the client's charts and coding rules, then the coders' corrections from each round.
  4. Try smaller.A final round tests smaller reasoning models and different representations of the chart, because the cheapest model that is accurate enough is the one that should run at volume.

In every round the client's coding experts check not only the codes but the reasoning. A right code for a wrong reason is a wrong code that has not failed yet.

How accurate it is

Accuracy is measured against the codes a coder verified. Across nine waves and 19,139 charts:

CodeMatched the verified codeCharts measured
Qualifying circumstances99.87%19,139
Physical-status modifier98.50%19,139
Anaesthesia code83.48%19,139
Mode of anaesthesia82.48%19,139
Procedure code73%100, in the latest wave

The pattern is the one the design expects. Where a code follows rules — whether a qualifying circumstance applies, which physical-status modifier the assessment implies — the engine is right almost every time. Where the code takes judgement, it is right about five times in six, and that gap is exactly why a coder sees every code. Procedure codes joined the measurement only in the latest wave, and are the hardest field of all.

The feedback loop shows most clearly in the mode of anaesthesia: 76.5 per cent in the first wave, and above 93 per cent in every wave since the fifth.

The data rules came first

The engagement started with where the data could go, before anything was built, and the same rules hold in production. Charts are stripped of identifiers before any model sees them. The work runs in the client's own cloud account, which their IT team controls, and what the engine learns stays in a vector store inside that account. Records sent to the model provider for processing are covered by terms under which the provider does not train on them.

A second specialty

The same approach has since been proven on internal medicine, where the input is different — progress notes arriving through an API rather than scanned charts — and so is the answer: up to twenty diagnosis codes, procedure codes each with its justification, place and type of service, dates of service, modifiers, and the mapping between procedures and the diagnoses that justify them. Bulk runs, accuracy checks and coder audits confirmed it can be coded the same way.

Client described rather than named, except where cleared.