The textbook way to build a credit model is to train on outcomes. Take historical loans, label each one by whether it was repaid, learn what predicts the difference.
This requires a history of loans with known outcomes. The digital lender we built the platform for did not have one. Worse, the loans needed to create one are exactly the ones nobody can responsibly make — lending to a large sample of applicants with no basis to assess them, then waiting a year to find out how badly it went.
It is a genuine cold start, and it is the reason new lenders begin with manual underwriting: credit officers, reading applications, making judgements.
The judgements are the dataset
Our approval models were trained on roughly eight thousand decisions previously made by hand by credit officers. Not on repayment outcomes. On the decisions themselves.
That reframes the objective in a way worth being explicit about, because it changes how the result is evaluated:
Predict repayment
Target: did this loan perform? Requires a mature book. Evaluated against realised defaults. Can beat human judgement, because it is learning from reality rather than from opinion.
Reproduce the underwriter
Target: what would a good credit officer decide? Requires only a manual history. Evaluated against held-out human decisions. Cannot beat the officers — that is not the goal.
The system is not trying to out-predict the credit team. It is trying to be the credit team, seven thousand times a day.
Once framed that way, the economics are obvious. A credit officer takes minutes per file and there are only so many of them. Reproducing their judgement at machine throughput is what converts a manual lending operation into a digital one — and it does so without waiting a year for a default history that does not exist yet.
What we actually inherited
We were cloning accumulated expertise, and that is a real asset. A credit officer's sense that a particular combination of employment tenure, existing obligations and bank balance volatility looks wrong is genuine information, built over thousands of files, and it is not written down anywhere else. Training on their decisions is the only practical way to extract it.
We were also cloning everything else they do.
This is not an argument against the approach. It is an argument for knowing what we built.
What checking actually involves
The instinct is to withhold the sensitive attributes and consider the matter handled. That does not work: correlated features reconstruct them readily, and a model that never saw a location can still act on one through employer, spending pattern or device.
Testing has to be on outcomes rather than on inputs, and for a judgement-trained model there is a specific advantage — there is a control group. The officers' own decisions are right there, so the question is not only whether the model treats segments differently, but whether it treats them differently from the way the officers did. Five checks follow from that:
- Disparate outcomes by segment, across protected attributes and their likely proxies — location, employer type, education, device tier, name-derived signals — whether or not any of them was a feature.
- Divergence from the officers, by segment. A model that agrees with human decisions ninety per cent of the time overall but only seventy per cent for one group has found something, and it is worth knowing what.
- Confidence by segment. Systematically lower confidence on a group usually means that group was thin in training, and thin segments are where cloning degrades fastest.
- Training set composition. Which segments were rare in the eight thousand? Rarity in the labels becomes unreliability in the model, and it is invisible in aggregate accuracy.
- Drift. The applicant population moves. A check that passed at launch is evidence about launch, not about now, so it has to recur.
None of this is exotic, and most of it is a few queries against decisions already held. The reason it often does not happen is not difficulty — it is that nobody owns it, and the model is performing well on the metric anyone is watching.
The unexpected benefit
Judgement-trained models are easier to get approved.
A risk committee asked to authorise automated lending has one central worry: that the machine will do something the institution would not have done. With an outcome-trained model, that worry is legitimate — the model has learned patterns from reality that may genuinely contradict institutional practice, and defending it means defending the maths.
With a judgement-trained model, the answer is structurally different. Its decisions are recognisably the decisions the committee's own officers were making, because that is exactly what it was fitted to. The conversation shifts from do we trust this model to do we trust our underwriting — a question they have already answered.
Pairing it with interpretable model families helps for the same reason. Decision trees for approve-or-decline give a path that can be read aloud, which matters when a declined applicant is entitled to a reason and a regulator is entitled to ask how it was produced.
What comes next, by design
Judgement-training is a bootstrap, not a destination. Every loan the automated system writes generates the outcome data that was missing at the start, and after a couple of cycles the lender has the dataset it originally wanted.
The transition needs planning rather than drifting, and it has a specific hazard: the outcome data is censored by the model's own decisions. Repayment is observed only for approved applicants, and they were approved because the judgement-trained model said so. Applications it declined have no outcome at all, so naively retraining on outcomes teaches the next model that the previous model's boundary was correct — including wherever it was wrong.
Mitigating that means deliberately preserving some signal outside the boundary: a small randomised approval band, or shadow-scoring declined applications against later observable behaviour. Neither is free. Both are much cheaper than discovering three model generations later that the model has been optimising inside a box it drew itself.
The generalisable part
When you cannot train on outcomes yet, expert decisions are a legitimate and underused label source. Any domain with skilled manual operators — underwriting, triage, moderation, quality inspection, dispatch — has a history of judgements sitting in a workflow system, and that history is a dataset. It is usually dismissed because it is not ground truth. It is not ground truth, but it is available now, and it encodes expertise that exists nowhere else.
Two conditions make it work. Know what you are cloning — audit for the patterns you did not intend to learn, because they will be there. And plan the exit — treat the judgement-trained model as the mechanism that generates your real dataset, and design for the censoring problem before it has three generations of history behind it.