Before a hospital or a practice can be paid for a procedure, somebody has to turn the clinical record into codes: what was done, why, under what kind of anaesthesia, for a patient in what condition. Payers pay against the codes, not the notes. The client does that coding for healthcare providers, and an experienced coder takes ten to fifteen minutes a chart.
That is the constraint the business runs on. Coders are skilled, scarce and expensive; volume arrives in batches; and a code that is wrong costs money in one direction or a compliance problem in the other.
We built an AI coding engine, starting with anaesthesia charts. It is live and integrated with the client's revenue-cycle systems. It reads the chart, assigns the codes, explains its reasoning and the alternatives it considered, raises alerts where the documentation is missing something the code depends on — and hands every code to a coder, who accepts or rejects it.
19,000+
Anaesthesia charts processed by the engine so far.About 75 seconds
Against ten to fifteen minutes by hand. The coder's job becomes checking the engine's argument rather than building it.5,000–7,000
Charts a day in bulk processing.What makes this hard
- The input is a scanned record, not a form. A single case runs to a dozen pages — consent, pre-operative assessment, the anaesthesia record itself — and parts of it are handwritten.
- The answer is several codes that depend on each other. The procedure code and up to eight ancillary codes, medical direction, the anaesthesia code, a physical-status modifier, the mode of anaesthesia, any qualifying circumstances, and up to ten diagnosis codes behind them.
- The rules move. The code sets are revised every year, and individual payers add rules of their own.
- The data is protected health information. Patient records cannot be sent anywhere casually, and every design choice starts from that.
- A confident wrong answer is the worst answer. A coder can work with "unclear". A plausible code that happens to be wrong gets billed.
The engine's most important output is not a code. It is the reason for the code, and the admission when the record does not contain one.
From chart to codes
A reasoning model rather than a plain language model was a measured choice: on documents like these it gave more accurate answers. The cost of that choice is that it needs clean text to reason over, which is why reading and extraction are separate steps in front of it rather than one model doing everything.
Which reasoning model was measured too. The engine runs on o4-mini, chosen after comparing it with larger models — GPT-5.1, o4 and Claude models among them. At thousands of charts a day, the smallest model that is accurate enough is the right one, and the comparison is what tells you which that is.
Showing the work
Every code comes back with two things a coder can check: the reasoning behind it, and the other codes the model considered and why it rejected them. That turns review from re-coding the chart into checking an argument — faster, and far easier to audit.
It also exposes where the model is weakest, which is the point.
When the record does not say
The instructive failure came early. For some procedures, the correct code depends on a detail — the material used, say — that the notes simply do not state. Asked to code such a chart, the model narrowed the answer to two codes and explained that the record did not settle it. Forced to choose one, it picked the more common option.
That is exactly what a good coder would not do. A good coder sends a query back to the clinician. The typical answer is not the answer for this patient, and billing it is a guess dressed as a decision.
So the engine is built to say so. Where the documentation is missing something a code depends on, it raises a missing-data alert rather than filling the gap with the most likely value — and the coder, not the model, decides what happens next.
Measured before it was trusted
The proof of concept was designed to prove the engine on the client's own history before it touched live work, in five rounds.
- Hold out a test set.A share of charts with verified coding is set aside from the start and never used for training, so accuracy is always measured on charts the model has not seen.
- Measure the untrained model first.The first round codes without any of the client's charts, so every later improvement is measured against what the model knew on its own — not against nothing.
- Train, then feed back.Later rounds add the client's charts and coding rules, then the coders' corrections from each round.
- Try smaller.A final round tests smaller reasoning models and different representations of the chart, because the cheapest model that is accurate enough is the one that should run at volume.
In every round the client's coding experts check not only the codes but the reasoning. A right code for a wrong reason is a wrong code that has not failed yet.
How accurate it is
Accuracy is measured against the codes a coder verified. Across nine waves and 19,139 charts:
| Code | Matched the verified code | Charts measured |
|---|---|---|
| Qualifying circumstances | 99.87% | 19,139 |
| Physical-status modifier | 98.50% | 19,139 |
| Anaesthesia code | 83.48% | 19,139 |
| Mode of anaesthesia | 82.48% | 19,139 |
| Procedure code | 73% | 100, in the latest wave |
The pattern is the one the design expects. Where a code follows rules — whether a qualifying circumstance applies, which physical-status modifier the assessment implies — the engine is right almost every time. Where the code takes judgement, it is right about five times in six, and that gap is exactly why a coder sees every code. Procedure codes joined the measurement only in the latest wave, and are the hardest field of all.
The feedback loop shows most clearly in the mode of anaesthesia: 76.5 per cent in the first wave, and above 93 per cent in every wave since the fifth.
The data rules came first
The engagement started with where the data could go, before anything was built, and the same rules hold in production. Charts are stripped of identifiers before any model sees them. The work runs in the client's own cloud account, which their IT team controls, and what the engine learns stays in a vector store inside that account. Records sent to the model provider for processing are covered by terms under which the provider does not train on them.
A second specialty
The same approach has since been proven on internal medicine, where the input is different — progress notes arriving through an API rather than scanned charts — and so is the answer: up to twenty diagnosis codes, procedure codes each with its justification, place and type of service, dates of service, modifiers, and the mapping between procedures and the diagnoses that justify them. Bulk runs, accuracy checks and coder audits confirmed it can be coded the same way.