The client lends to salaried employees, mostly early in their careers, in amounts up to about ₹25 lakh. Many are borrowing formally for the first time. Their credit bureau file is thin or empty — not because they are bad risks, but because nobody has lent to them before.
A bank handles that with a branch visit, documents and a human underwriter, over days. That process costs more than the interest on a small loan earns, so it does not scale down. The entire business rests on making a sound credit decision in minutes, from data the applicant can supply through a phone, for a loan whose median size is around two thousand dollars.
Get it slightly wrong in one direction and you decline good customers. Slightly wrong in the other and the losses exceed the margin. There is no comfortable middle.
What makes this hard
- The decision has no single source of truth. No one input is sufficient, several are missing for any given applicant, and the ones that exist disagree with each other.
- Latency is a product feature, not a target. Fifteen external checks have to resolve inside a window the applicant will tolerate — and several of those services are slow, rate-limited, or occasionally down.
- Fraud adapts. Static rules catch last quarter's patterns. The system has to flag applications that are merely unusual, not merely non-compliant.
- Being wrong is expensive in both directions, and the feedback loop is months long. You learn whether today's decision was right when the borrower does or does not repay next year.
- It is regulated infrastructure. Lending partners, bureau reporting, mandate execution and data protection all carry obligations that do not bend to a release schedule.
The hard part of automated lending is not the model. It is deciding, for each applicant, which evidence to trust.
An origination pipeline, end to end
Application through disbursal, then the whole life of the loan afterwards.
Not one model — four analysers and an arbiter
The credit decision is decomposed rather than monolithic. Four independent analysers each read a different kind of evidence, produce their own view, and feed a final calculation that combines them into an internal score.
Credit history
Score, existing obligations, repayment record. Authoritative when it exists — and absent or near-empty for a large share of first-time borrowers.Statement analysis
Salary credits, expense patterns, balance behaviour, existing instalments. Establishes real disposable income rather than declared income.Device signals
What the handset and its applications indicate about stability and behaviour, for applicants with no formal record at all.Transactional messages
Salary credits, instalment debits and bank alerts as corroborating evidence — often the only proof of income for a first-time borrower.Decomposition is the design decision worth defending. The analysers fail independently, and that is the point. A thin-file applicant has no bureau record but does have bank statements. Someone in their first job has neither, but has salary messages. A single model over a wide sparse feature vector would degrade quietly across all of them; four analysers degrade visibly, one at a time, and the arbiter knows which evidence it is missing.
Two jobs, two model families
Approval and pricing are different problems and get different treatment. Decision trees handle approve-or-decline, branching on identity, bureau signals, banking behaviour and employment — a discrete outcome from interpretable splits, which matters when a declined applicant is entitled to a reason. Multivariate regression sets the interest rate, because a rate is a continuous quantity and should come from a continuous model. A third tree routes each approved application to the financing partner best suited to fund it.
The training set is the interesting part
The approval models were trained on roughly eight thousand decisions previously made by hand by credit officers. Not on repayment outcomes — on human judgements.
That framing changes the objective. The system is not trying to out-predict the credit team; it is trying to reproduce their judgement at seven thousand applications a day. It inherits their accumulated sense of which combinations look wrong, and it inherits it in a form that runs in milliseconds. Repayment data then refines the policy over time, but the cold-start problem — how do you underwrite before you have a default history — is solved by learning from the people who were already doing it well.
It also makes the system explainable to the risk committee that has to sign it off, because its decisions are recognisably the decisions the committee's own officers were making.
Proving who someone is, from a phone, in seconds
Onboarding without a branch visit means the identity check has to be as rigorous as a counter clerk's and considerably faster. Two computer-vision models do the work before any government check is called.
An object-detection model reads the identity document — locating and extracting fields from a photograph taken in poor light, at an angle, on a mid-range handset, which is the realistic input rather than a flatbed scan. A separate face model matches the applicant's selfie against the photograph on that document, with a managed vision service as a second opinion.
Only then do the authoritative checks run against government identity and employment records. Running the cheap local checks first means an obviously mismatched application is rejected in seconds without consuming a paid third-party call — which matters at seven thousand applications a day.
How did four engineers build all of this?
Forty-odd repositories cover this platform: customer apps on Android, iOS and Flutter, sales and partner portals, collections, operations consoles, credit-policy dashboards, KYC, analytics, test automation. That is a lot of surface for four people, and the answer is not heroics.
Build the decision, integrate everything else
Identity verification, bank statement aggregation, bureau data, mandate execution and disbursal are all bought rather than built — they are commodity, regulated, or both, and building them would have consumed the team without differentiating the product.
What stayed in-house is the part that is genuinely theirs: the credit decision, the scoring engine, and the orchestration that holds the pipeline together. That discipline about scope is the reason the headcount is four rather than forty.
No big design up front
The architecture was designed for what the lender needed at launch and what it could clearly see coming, and no further. No months spent designing for every possible future, and no platform built ahead of the product. It grew as the business did, on cloud services rather than self-managed infrastructure, so that four people were writing lending logic instead of operating databases.
The model-serving choice is a small example that generalises. The credit engine runs as a function rather than a standing service: no cluster to keep warm, no capacity planning for a workload that arrives in daytime bursts, and nobody on the team owning an operational surface that adds nothing to the credit decision.
Built for an external assessor to look at
The platform holds identity documents, bank statements and salary data for hundreds of thousands of people. It has passed independent external vulnerability assessment, and the architecture was written to be examined rather than explained away.
| Control | Implementation |
|---|---|
| Data at rest | AES-256 across tables, indices, replicas, backups, caches, file storage and analytics lakes, with keys in a managed key service |
| Data in motion | TLS 1.3 on all API traffic |
| Key management | Hardware security modules validated under FIPS 140-2; keys centralised, rotated and audited |
| On the handset | Transient store-and-forward buffer; minimal persistence in an app-sandboxed local database |
| Credentials | Short-lived limited-privilege tokens — real user credentials never traverse the client-server boundary |
| Authorisation | Role-based control over features, attribute-based control over which accounts a user can reach |
| Audit | Encrypted event log covering both management and data events |
The authorisation split is worth a sentence, because it is the one most often collapsed. Roles govern what kind of thing you may do; attributes govern which instances you may do it to. Lending operations need both — a collections agent has a role, but they should only reach the accounts assigned to them.
Three shapes of lending engagement
A small team owns the product end to end, from credit models to mobile apps, and hands over a running business.
Embedded development pods on an established lender's origination and management systems, with engineering management inside the pod — including security remediation across shared libraries.
Current work on agentic AI for mortgage underwriting in North America: an orchestrator coordinating worker agents across the approval workflow.