Two systems sit on the critical path of every transaction through the UPI and IMPS switch we build and run for an ATM operator. The national scheme on one side, the sponsor bank's core banking system on the other. Neither is ours. Neither can be summoned when a developer wants to reproduce something.
The scheme's certification environment is a shared, scheduled resource — we book time in it. The bank's core banking system is available during integration windows negotiated between two organisations. We cannot load-test either. We certainly cannot ask either to fail in a specific way at a specific moment because we are debugging.
This is not unusual in enterprise integration. What is unusual is how often teams respond by simply testing less, and calling the result a dependency.
Build the other side
We built both counterparties: a stub of the scheme's request flow, and a separate core banking simulator. Not as test fixtures inside a test suite — as first-class code, in their own repositories, maintained like the product.
That distinction sounds like bookkeeping and is not. A fixture is something one developer wrote to get a test passing, and it rots. A simulator that is versioned, reviewed and owned is something the whole team can rely on to behave like the real thing — which is the only property that makes it useful.
Work stops waiting
Development against the simulated counterparties continues while certification slots are scheduled around it. The booking calendar stops being the critical path.Failures become reproducible
Timeouts, late callbacks, duplicate responses, acknowledgement-without-outcome. Provoked on demand, in a loop, in CI.Certification verifies
By the time the review rounds ran, the awkward paths had been exercised hundreds of times. Certification confirmed behaviour rather than discovering it.The second effect is the whole point
It is tempting to justify simulators on availability — we could not get access, so we faked it. That undersells them, and it leads to building the wrong thing.
The failure modes that matter most in an asynchronous payments protocol are the ones nobody can provoke against a live system, even with unlimited access:
- The outcome callback that never arrives.
- The callback that arrives twice.
- The callback that arrives after our process restarted.
- The acknowledgement followed by silence, where the debit may or may not have happened.
- The response that arrives out of order relative to another transaction.
We cannot ask a national payment scheme to drop a callback so we can see what our code does. But these are precisely the conditions under which a payments system loses money, and precisely the conditions under which a customer is charged for something they did not receive.
A simulator that only does the happy path is worse than no simulator, because it manufactures confidence about the part that was never in doubt.
So our design goal was not fidelity to the specification. It was fidelity to the specification's failure modes. The stub has to be able to be wrong in the specific ways the real thing is wrong — late, twice, never, out of order — and those behaviours have to be switchable from a test.
What it costs, and what it does not
The obvious objection is that we now maintain a second implementation of somebody else's system, and it will drift from the real one.
Both halves of that are true and neither is fatal, because we are not reimplementing the scheme. We implement the parts of its protocol behaviour that our code reacts to — the message exchange shape, the timing, the failure modes. That is a much smaller surface than the scheme's actual functionality, and it changes far more slowly.
We manage drift the way any contract is managed: certification rounds against the real environment are the reconciliation point. When the real thing surprises us, the surprise goes into the simulator, and it is never a surprise again. The simulator accumulates institutional knowledge of the counterparty's edge cases — which is a genuinely valuable artefact, and it is the thing that would otherwise live only in the memory of whoever was on the call.
The generalisable part
When a dependency is on your critical path and outside your control, the instinctive question is how do we get more access? — more environment time, a better sandbox, a friendlier integration window.
That is usually the wrong question, because access is negotiated with someone whose priorities are not yours, and because more access to a well-behaved system still will not show you what happens when it misbehaves.
The better question is: what do we need to be able to reproduce? Enumerate the behaviours your code must survive, and then build the smallest thing that can produce them on demand. Usually that is far less work than it sounds, because the list is shorter than the counterparty's feature set and consists almost entirely of failure.