The internal AI platform we built for a lending-technology company — a team of three — runs the company's finance review queues. Staged accounting transactions arrive; each needs a decision only the business can make, such as which client a cost served; a person confirms it.
The obvious way to use a model here is to hand it each transaction with everything known about it and ask for the whole answer. The first version of the queue did roughly that: batches of transactions to a capable model, a rich schema for the response — category, client, treatment, revenue recognition, budget impact, confidence, reasoning — and a human to confirm. It went into use. It also made the model the first thing consulted about every line, including lines whose answer was never in doubt.
When we rebuilt the pipeline, we inverted it.
The chain
For each line, the rebuilt pipeline asks a sequence of questions, ordered by how much each kind of answer can be trusted, and stops at the first that answers.
The reasoning for the order is recorded in the module itself: running the model over a line a rule already answers introduces a chance of disagreeing with a certainty, and there is no version of that trade worth making.
A model consulted about a question that already has an answer can only add a way to be wrong.
There are quieter benefits. The first three steps are database lookups; they cost nothing and run in milliseconds. Every confirmation a reviewer makes becomes a step-two answer the next time that vendor appears, so the model's share of the work shrinks as the team uses the system. And the model at the end of the chain is a small, fast one, because by the time a line reaches it, the question it is asked is narrow.
Starve the input
What the model sees is deliberately limited: the vendor name, the account, the class, the line's description, and the list of clients it may choose from. No amounts, no dates, no identifiers, no tax, no totals.
This is not about secrecy. It is the company's own data. It is about keeping the model's reasons honest. A model that can see an amount will use it, and "this is a large bill, so it is probably the large client" is a correlation, not a reason. Starving the input is what keeps the reason an explanation rather than a rationalisation.
The candidate list matters as much as the text. It is exactly the set of clients the reviewer's own picker offers, with archived clients excluded. A model choosing from a different list would propose clients a person cannot confirm — and a downstream check that tests whether a value is present, rather than whether it is one of the allowed values, would let it through.
Narrow the output
The model returns one thing, through a forced tool call with a fixed shape: where the cost goes, which client, a confidence of high, medium or low, and a reason — which must quote the text it relied on.
It cannot express an account, a treatment, a split between clients or a period. Those are either facts, which the model never sets, or decisions that belong to a person with an editor in front of them.
Validate every source, not just the model
Every proposer's output passes the same validation — the rules and the learned mappings included. A rule that returns a malformed proposal is the same defect as a model that does, and validating only the source that looks untrustworthy is how the trusted one gets away with it.
Validation also throws rather than returning nothing. A bug that quietly produced an empty result would look identical to a line nobody could answer, and would sit in the queue looking like a hard case rather than a broken rule.
Failing last is allowed
The unattended model is off by default. The first version had removed bulk AI on a timer precisely to stop automatic spend on every sync, and wiring a model into a thirty-minute schedule without a decision behind it would have quietly undone that. With it off, the first three steps still run, and cost nothing.
When it is switched on, it runs under a cap per cycle, inside the per-user budgets, so a first-run backlog cannot spend the day's allowance and starve the person waiting on an interactive request. And when the model cannot answer — the budget is spent, or the response is malformed — the line is simply left unproposed, for a person.
That outcome is only acceptable because the model is last. If it were first, its failure would be the pipeline's failure. At the end of the chain, a model that declines to answer leaves a person exactly where they would have been without it.
The generalisable part
Order your sources by how much you can trust them, and let the first answer win. Facts before rules, rules before past human decisions, all of them before inference. A model belongs at the end of that chain, where it handles only what nothing more reliable could — and where its share of the work shrinks as people confirm its proposals.
Give the model the least it needs, not the most you have. Every extra field is an invitation to reason from correlation, and you will not be able to tell from the output that it did. Choose the candidate set to be exactly what a human could confirm.
Make the output unable to say what you do not want it to decide. A schema is a boundary. If the model must never set a treatment, the response should have nowhere to put one.
And design the model's failure to be a no-op. If the model cannot answer, the right result is the state you would have been in without it — a question left for a person, not a guess, and not an error that stops everything else.