Engineering notesNo. 04ReliabilityAll notes

Idempotency by annotation, and the race that can undo it

Methods in the UPI and IMPS switch we run declare idempotency with an annotation, and one aspect enforces it, so coverage is a grep away. The trap is the obvious check-then-act enforcement, which lets duplicate requests through; a uniqueness constraint claimed first does not.

IdempotencyConcurrencyAOPPaymentsAPI design

Every payments integration has the same requirement written somewhere: operations must be idempotent. Sending the same instruction twice must not move money twice.

The usual way this gets implemented is as a convention. Somebody writes it in an onboarding document, a senior engineer mentions it in review, and each handler grows its own duplicate check. It works, for a while, in the way that all conventions work — right up until a new joiner writes the fourteenth handler and nobody notices what is missing.

A guarantee that depends on everyone remembering it is not a guarantee. It is a statistic.

Make it declarable

The alternative is to make idempotency something a method declares, and something the framework enforces. In the UPI and IMPS switch we build and run for an ATM operator, that is a marker annotation and an aspect around it:

The declaration

@Idempotent

An empty marker annotation, retained at runtime, targeting methods. It carries no configuration and no logic. Its entire job is to be visible.

The enforcement

One aspect

Advice around every annotated method, binding the incoming request. It extracts the request identifier, rejects the call if it is absent, and refuses the work if that identifier has been seen before.

Three properties follow, and the third is the one worth the effort.

  • The handler stays clean. No duplicate-checking code interleaved with business logic, so the business logic is readable as business logic.
  • The policy has one home. A change to how duplicates are detected is made once, not in fourteen places with three subtly different behaviours.
  • The guarantee becomes enumerable. Every method that claims to be idempotent can be listed with a grep, and the list diffed against the set of operations that ought to be. Coverage becomes a question with an answer rather than a matter of confidence.

That last point is what separates this from a tidier way to write the same conditional. A convention cannot be audited. An annotation can, and the audit takes a second.

The trap in the middle

Here is where a straightforward implementation goes wrong, and it goes wrong in the way concurrency bugs usually do — invisibly, under load, in production.

The obvious shape of the aspect is:

Check-then-act — the shape that looks correct
1ReadHas this request identifier been seen before?
2DecideNo. Proceed.
3ExecuteRun the handler. Money moves.
4WriteRecord the identifier so the next one is rejected.

Read, decide, act, record. It is correct in a single-threaded trace, and it is wrong the moment two copies of the same request arrive close together — which is exactly what happens when a client times out and a retry lands beside the original, or when two instances pick up duplicates from a queue.

Widening the window makes this worse in an obvious way, but narrowing it does not make the design correct. There is no implementation of check-then-act that is safe, because the gap between the check and the act is where the second caller lives.

Let the database arbitrate

The fix is not a lock, a cache or a distributed coordination service. It is to stop asking a question and start making a claim.

Insert the request identifier with a uniqueness constraint on it, before doing the work. One of the two concurrent callers succeeds. The other gets a constraint violation, which is exactly the signal needed — and it translates directly into the duplicate rejection.

Check, then actRacy

Two operations with a gap between them. The gap is the bug. No amount of narrowing removes it, and it fails under precisely the conditions idempotency exists to handle.

Claim, then actCorrect

One atomic operation decides the winner. The uniqueness constraint is doing the mutual exclusion, and the database was always going to be better at that than application code.

Two details matter in the implementation. The claim and the work should share a transaction boundary, or an identifier can end up claimed for work that then fails — turning a retriable failure into a permanent one. And the claim record should store the outcome once it is known, so that a genuine duplicate can be answered with the original result rather than merely refused. A caller who retries because they never saw the response wants the answer, not an error.

The generalisable part

Two things, one narrow and one broad.

Narrow: any time you find yourself reading state to decide whether to write it, check whether the storage layer can make that decision atomically instead. Unique constraints, conditional writes and compare-and-set exist precisely so application code does not have to invent mutual exclusion. Application-level check-then-act is one of the most common concurrency bugs in ordinary business software, and it is almost always avoidable rather than merely mitigable.

Broad: when a property must hold across a whole codebase, prefer a mechanism that makes the property visible over one that makes it implicit. A convention degrades silently with every new joiner, and its coverage cannot be measured. A declaration can be listed, counted and diffed against what ought to be there. The declaration does not make the code more correct on its own — the enforcement does that — but it makes incorrectness findable, which over a codebase's life is worth more.

Engineering notes — No. 04 · 2026-10-01