Engineering notesNo. 12Applied AIAll notes

An LLM for medical coding that says what is missing

Made to choose a code the chart could not support, the model in the medical coding engine we built picked the most common one. So the engine flags the missing detail instead of guessing, and a coder asks the clinician. It has coded more than 19,000 anaesthesia charts.

Applied AILLMsHealthcareUncertainty

The AI coding engine we built for a US healthcare revenue-cycle company — three engineers and a business analyst — reads anaesthesia charts and assigns the billing codes, with its reasoning, for a coder to accept or reject. It has coded more than nineteen thousand charts.

The most useful thing we learned building it came from one chart, early in the proof of concept.

Two questions, two answers

The chart described a procedure whose correct code depends on a detail — which of two ways it was performed — that the notes did not state. We asked the model to code it.

Its first answer was exactly right. It narrowed the choice to two codes, explained that the code turns on that detail, noted that the record did not settle it, and said a coder would need to query the clinician before assigning one.

Then we asked it to choose one anyway, autonomously. It picked the more common of the two, on the grounds that most such procedures are done that way — while still noting that a coder should confirm.

The second answer is the one a pipeline would have kept. A field that must be filled will be filled, and the model filled it with the base rate: the answer that is right for most patients, and wrong for this one whenever this one is the exception.

Why this happens

Take an invented case to see the mechanism plainly. Suppose a procedure is coded differently depending on whether it was done on one side of the body or both, and a note records the procedure without saying which. Nothing in the note makes either answer more likely for this patient. The only evidence available is the population: which is more common in general.

Forced to choose, a model does the statistically sensible thing and picks the common case. Nothing about its output shows that it did. The code arrives in the same field, in the same format, with the same confidence of presentation as a code the record fully supported. Downstream, a coin-flip and a certainty look identical.

A forced choice turns "the record does not say" into "the most common answer", and hides the conversion.

This is not a failing of one model. It is what the question asked for. If the output schema has one place for a code and no place for "cannot be determined from this record", every uncertainty will be resolved silently, in the direction of the average.

Give it somewhere to say so

So the engine's output has a place for what is missing. Where the documentation lacks something a code depends on, it raises a missing-data alert rather than filling the gap with the likeliest value. Every code also comes back with its reasoning and the alternatives the model considered, so a coder can see when a choice was close.

And the decision stays with a person. The coder sees each code, can accept or reject it, and — for an alert — can query the clinician, which is what a good coder would have done without any AI involved. The engine's job is not to remove that query. It is to make sure the query gets asked.

Where the flags matter

The engine's accuracy, measured against coders' verified codes, shows why. On the fields that follow rules — qualifying circumstances, the physical-status modifier — it matches almost every time: 99.87 and 98.50 per cent across 19,139 charts. On the fields that take judgement — the anaesthesia code, the mode of anaesthesia — it matches about five times in six.

Those judgement fields are where records are most often ambiguous, and where a forced answer would do the most damage. A system that hid its uncertainty would present them with the same confidence as the rule-bound fields. Ours does not.

The same principle, elsewhere in our work

We keep arriving at this in systems that have nothing to do with healthcare.

Month-end finance

Ask for the split

On the internal AI platform we built, costs shared across clients are split at month-end by each client's revenue share. When no client had revenue in the window, the engine asks a person for the split instead of guessing one.
Finance queue

Leave it unproposed

When the model at the end of the allocation chain cannot answer — its budget is spent or its response is malformed — the line is left for a reviewer rather than filled in.
Credit decisions

Decline to decide

The arbiter over four credit analysers knows which evidence is missing for each applicant, and can decline to decide or route to a person rather than produce a confident score from too little.

The generalisable part

Give your model a first-class way to say "not determinable from this input". Not a low confidence score on a forced answer — an explicit output that means the input does not contain what the answer depends on. If your schema has no place for it, every uncertainty will be resolved silently in the direction of the average.

Count abstentions separately from errors. A correct "the record does not say" is not a miss, and a confident wrong answer is worse than either. If your accuracy figure treats them the same, it will reward the model for guessing.

Route the abstention to the person who can get the missing fact. An alert is only useful if it reaches someone who can act on it — here, a coder who can query the clinician. Design that route before you design the model.

And be suspicious of any required field downstream of a model. A field that must be filled will be filled, and what fills it when the evidence is absent is the base rate, presented as a fact.

Engineering notes — No. 12 · 2026-10-01