The consumer lending platform we built for a digital lender in India — four engineers at launch, now deciding seven thousand applications a day — has about five minutes to decide whether to lend to somebody it has never met. The evidence available is whatever can be gathered through a phone in that window: a credit bureau record, bank statements, signals from the device itself, and the transactional messages the bank sends to it.
The obvious architecture is one model. Concatenate every feature from every source into a vector, train on historical decisions, get a score. It is simpler to build, simpler to deploy, and it is what most teams reach for.
We built four independent analysers feeding an arbiter instead. The reason is not accuracy.
Credit history
Score, existing obligations, repayment record. Authoritative — where it exists.Statements
Salary credits, spending, balances, existing instalments. Real disposable income rather than declared income.App signals
What the handset indicates about stability, for applicants with no formal record at all.Transactional messages
Salary credits and instalment debits as corroboration — often the only evidence of income a first-time borrower has.The inputs are not merely noisy. They are absent.
This is the distinction that decides the architecture, and it gets blurred constantly.
A noisy feature is present and unreliable. Models handle that well; it is what regularisation is for. An absent feature is a different problem, and in this population absence is not an edge case — it is the norm, and it is structured rather than random:
Now consider what one model over a wide, sparse vector does with these. It produces a score for every one of them. It does not say that applicant B was scored almost entirely on message parsing, or that applicant C's statement features were actively misleading. The degradation is real, continuous and silent.
A single model over sparse inputs does not fail. It quietly becomes less right, and it reports that with exactly the same confidence as when it was right.
What decomposition buys
Four analysers, each with its own feature extraction, each producing its own view. An arbiter combines them into an internal score.
The property this buys is not a better score on the average applicant. It is that the arbiter knows which evidence it is missing, and can behave differently as a result. It can weight what is present, decline to decide where too little is present, and route an unusual evidence profile to a human instead of guessing with a confident number.
Three further consequences follow from the same structure:
- Failures are attributable. When the bureau is unreachable, one analyser is unavailable — not the decision. The platform can degrade deliberately instead of discovering later that a whole day of decisions was made on a silently-empty feature block.
- Components evolve separately. Statement parsing changes when aggregator formats change. Message parsing changes when banks reword alerts. These have nothing to do with each other and should not share a release, a model version, or a regression suite.
- The decision is explainable by construction. A risk committee can ask which evidence drove an outcome and get an answer with a structure, rather than a feature-importance chart over two hundred columns.
The cost, honestly
Decomposition is not free and the trade is real.
We gave up the cross-source interactions a joint model could learn — the subtle relationship between a spending pattern and a device signal that no single analyser sees. For some populations that is a genuine loss of predictive power; where applicants all have complete data, the joint model is the better answer.
We also took on an arbiter, which is a second thing to design, calibrate and defend. Combining four views is its own modelling problem, and doing it badly can cost more than the decomposition gained.
The trade is worth making when inputs have independent availability — when the reason a feature is missing is a fact about the applicant rather than a glitch. That is exactly the case in first-time borrowing, and it is why the architecture follows the population rather than the maths.
The applicant profiles above are illustrative of the population rather than documented segments.
The generalisable part
When you are choosing between one model over everything and several models over parts, the question is usually framed as accuracy. It is more useful to ask about availability structure.
If your inputs fail independently, decompose along the failure boundaries. You will trade some cross-feature signal for the ability to know, at decision time, what you are missing — and in any domain where a wrong answer is expensive and a deferred answer is merely inconvenient, that knowledge is worth more than the marginal accuracy.
The broader version: observability of the inputs is a modelling decision, not an operational afterthought. An architecture that can say "I decided this with three of four sources, and the missing one was the authoritative one" is more useful in production than a marginally better score that cannot.