Engineering notesNo. 03PaymentsAll notes

A UPI payment is two conversations

UPI acknowledges a payment request on one connection and reports the outcome later by calling our switch. How that shaped the switch we run for an ATM operator, and why a missing response calls for a status check, not a retry that could debit someone twice.

PaymentsAsynchronous protocolsIdempotencyIntegration designFailure handling

Most systems answer on the connection they were called on. Send a request, wait, get a result. The mental model is a function call with network latency, and almost every HTTP client, retry library and circuit breaker assumes it.

Real-time payment schemes do not work this way. We build and run the UPI and IMPS switch for an ATM operator — a team of three built it, reconciliation included — and India's UPI is the scheme we integrate with most; the same shape appears in national instant-payment rails elsewhere. Our switch sends a request and gets back an acknowledgement that the scheme received it. Then the connection closes, and some time later the scheme calls us with what actually happened.

One payment. Two independent HTTP exchanges, in opposite directions, minutes apart.

One payment, two conversations
1Switch → schemePay request. Signed, over mutual TLS, carrying an identifier we generate and will need again later.
2Scheme → switchAcknowledgement. "Received." Nothing about whether money moved. The connection closes here.
⋯The indeterminate windowThe debit may have happened. It may not. There is no way to know from anything we hold. The customer is watching a spinner.
3Scheme → switchPay response — inbound, on our endpoint. The actual outcome, arriving as a fresh request from the scheme, which we must correlate back to the original.
4Switch → schemeStatus check. Only when the response never came. A different question, not a repeat of the first one.

The acknowledgement is the most dangerous message in the protocol

It arrives on the connection where a result would normally be. It has a success status. Every instinct built from ordinary API work says: the call worked.

It means the scheme has the message. That is all. The payment may subsequently succeed, fail, or be rejected for a reason that has not been determined yet.

We are not a client. We are a server.

The second conversation is inbound. The scheme makes a request to an endpoint we expose, and expects us to be there.

That inverts the usual integration posture, and the consequences are not small. Our availability becomes part of the scheme's success path. Our certificates and network configuration have to admit its traffic. We need an endpoint that is reachable, authenticated, and fast enough not to time out on the scheme's side — during deploys, during restarts, at three in the morning.

It also means outcomes arrive at an instance that may not be the one that sent the request, minutes later, possibly after a rolling restart. Correlation cannot live in memory. The identifier we generate on the way out has to be durable, indexed, and matched on the way back in — and that requirement shaped our data model long before it shaped the code.

The retry that isn't a retry

This is the part that matters well beyond payments.

The response does not always come. Networks fail in both directions, and a scheme under load may simply not call back. That leaves a transaction in an unknown state and a customer waiting.

The reflex — the one every HTTP client library encourages — is to retry the request. In payments, that is how somebody gets debited twice. The original request may have been processed perfectly; the only thing that failed was the outcome reaching us. Sending it again asks for a second payment.

The correct move is not to repeat the question. It is to ask a different one.

The schemes provide a status-check verb for exactly this, and the distinction is precise:

pay

Move this money

Has side effects. Unsafe to repeat. Sending it twice may move money twice.

status check

What happened to transaction X?

No side effects. Safe to repeat indefinitely. Returns the same answer every time.

So our recovery path for a missing response is not a retry loop on the payment — it is a query loop on the status. The switch can poll that as often as it likes, for as long as it likes, without any risk of moving money.

Idempotency keys are the other half of the same defence, and they are not a substitute for this. An idempotency key protects against the same request arriving twice. A status check removes the need to send it twice at all. Both are needed: the key makes duplicates harmless, and the status verb means they are rarely created.

The package structure ends up as the protocol

When outcomes arrive independently of requests, the two halves are genuinely separate code paths — different entry points, different authentication direction, different failure modes, different tests. Trying to hide that behind a single synchronous-looking function is a leaky abstraction that leaks at the worst moments.

So our switch has a module for each half of each verb. The pay request and the pay response are siblings, not caller and return value. Status check and status response likewise. Heartbeat and its response. Address validation, account-provider lookup, and the dispute-resolution verbs all sit alongside them as first-class modules rather than helper functions hanging off a payment client.

The directory listing reads as a description of the protocol. That is not an accident of organisation — it is the structure the protocol forces, and fighting it produces code that is harder to reason about, not easier.

One envelope, several transactions

A related trap: the pay verb is not one operation. In our implementation it fans out into distinct flows — an ordinary debit, a refund credit, a dispute adjustment, and a mandate update. Same message envelope, genuinely different transactions.

It is tempting to treat them as one handler with a type field and a switch statement. Do that and the dispute-adjustment path inherits the debit path's assumptions, which is a category of bug that surfaces during a regulator-facing reconciliation rather than in a test.

The generalisable part

Payments make this vivid because the money is real, but the pattern is not specific to them.

Whenever an operation has external side effects, a timeout is not a signal to repeat it. It is a signal that you do not know the state — and unknown state is resolved by asking, not by acting. If the system you integrate with offers a way to query the outcome of an operation by identifier, that is your recovery path, and it is worth building even when retrying looks simpler.

If it does not offer one, that is worth knowing early, because it means every timeout leaves you with a decision that has no safe answer.

And: an acknowledgement is not a result. Any protocol that separates "I have your message" from "here is what happened" will eventually be integrated by somebody who conflates the two. Usually at three in the morning, usually for a customer who has definitely been charged.

Engineering notes — No. 03 · 2026-10-01