System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

Payment System deep dive: double-payment prevention and consistency across services


The payment flow from the first lesson wrote an idempotency key before calling the provider and published events through an outbox. This lesson explains why each of those details is exactly where it is. The first half is the central mechanism of the section, and the thing interviewers most reliably probe: stopping a double charge.

The idempotency key, done properlyClientgenerates a keyINSERTkey, same txnDuplicatekey rejectedReturnstored responseKey expiresafter 24 hInsert first, inside the charge transaction, or two racing retries both pass.
Checking whether a key exists before inserting it is exactly the read-then-write race that charges the customer twice.

The three ways to charge twice

The customer retries. Double-click, browser refresh on the confirmation page, or tapping Pay again after a slow response.

The client retries automatically. A mobile application with a 10-second timeout and an automatic retry, against a provider call that took 12 seconds and succeeded.

Your own service retries. A queue redelivers a payment task after a worker died mid-flight — the at-least-once delivery of the message queue's delivery semantics, arriving at money.

All three produce the same shape: two requests expressing one intention. The fix has to recognise the intention, not the request.

The idempotency key

The client generates a unique key — a UUID is fine — for one payment attempt, before the first send, and reuses it on every retry of that same attempt.

The server stores it with a unique constraint:

SQL
CREATE TABLE idempotency_key (  key             TEXT PRIMARY KEY,  request_hash    TEXT NOT NULL,  payment_id      BIGINT,  response_status INT,  response_body   JSONB,  created_at      TIMESTAMPTZ NOT NULL DEFAULT now());

The rule: insert the key as the first statement of the transaction that creates the payment. Not after the provider call. Not in a separate transaction. First.

Why "first, in the same transaction" matters

Consider inserting the key after a successful provider call instead. Two requests arrive together. Neither finds a key. Both call the provider. Both charge. The key stops the third request and not the second — which is the case that actually happens.

Now consider inserting it first, in the transaction. Request A inserts the key and holds the uncommitted unique-index entry. Request B attempts the same insert and blocks on it. When A commits, B's insert fails with a duplicate-key error, and B reads A's stored response and returns it. When A rolls back, B's insert succeeds and B proceeds as the only attempt. Both outcomes are correct, including the concurrent case.

This is the same ordering point as the hotel booking lesson, and it is worth stating precisely because most candidates describe idempotency keys without describing when the key is written — which is the only part that determines whether they work.

Storing and returning the original response

A retry must return what the first attempt returned, not a fresh result. Store the status code and response body against the key. A retry after success returns the original success with the original payment identifier, so the client's state converges rather than diverging.

If the same key arrives with a different request body — same key, different amount — that is a client bug or an attack. Return a conflict error and do not process. Storing a hash of the request is what makes this detectable.

Scope and lifetime

Scope the key per merchant or per API credential, not globally, so two clients cannot collide and one client cannot probe another's keys.

Lifetime: keep keys for at least as long as any retry could plausibly arrive, and longer for audit. Twenty-four hours is a common minimum; keeping them for the full retention period of the payment is safer and cheap at 10 million rows a day. Deleting keys too early reopens the duplicate window in exactly the situation — a delayed retry after an outage — when it is most likely to be exercised.

Consistency across services

Idempotency keeps one service honest. But the payment service, the order service, and the ledger cannot share a transaction. Neither can you and the provider. So what holds the system together?

Two ways to keep three services agreedDistributed transaction• Two-phase commit across services• One slow participant blocks the rest• No card network will ever join oneSaga with compensation• Each step commits locally, at once• Failure runs a compensating step• A refund is not a rollback
Compensation in payments is visible to the customer: the money moved and moved back, and both entries stay in the ledger.

Why a distributed transaction is the wrong answer

Two-phase commit gives atomicity across participants and is rejected here for reasons that are sharper in payments than anywhere else.

External participants cannot enrol. A payment provider will not join your two-phase commit. This alone settles it, because the provider is the participant that matters.

It blocks on coordinator failure. Between prepare and commit, participants hold locks. If the coordinator dies, they hold them until a human resolves it — during which payments stop. For a system whose priority is correctness, a stuck-but-consistent state sounds tolerable, and in practice it means an outage with locks held across the money path.

Availability multiplies down. Four participants at 99.95% give roughly 99.8%, which is about 17 hours of downtime a year created purely by coupling.

The saga, and why compensation in payments is not rollback

A saga is a sequence of local transactions, each with a compensating action.

For a payment: reserve inventory, authorise payment, create the order, capture payment. If creating the order fails, compensate by voiding the authorisation and releasing inventory.

The critical difference from a database rollback: a compensating transaction is a new, visible, forward action. Voiding an authorisation before capture is close to invisible to the customer. Refunding a captured payment is not — the customer sees a charge and then a credit, possibly days apart, possibly on a statement, and possibly with a foreign-exchange difference between the two.

That has a design consequence worth stating: order the saga so that the irreversible step is last. Authorise early, capture late. An authorisation is cheap to void; a capture is expensive to undo. Sequencing by reversibility is a real technique, and naming it distinguishes a candidate who has thought about sagas from one who has read about them.

Compensations must also be idempotent, because the saga coordinator retries them, and they must be recorded in the ledger as their own entries — never as a deletion of the original.

The transactional outbox

Step 2 of any saga does two things: change local state and tell someone. Doing them separately is unsafe.

  • Commit, then publish: the process dies in between and the event is lost. The payment is captured and the order is never marked paid.
  • Publish, then commit: the transaction rolls back and the event is already out. Downstream systems act on a payment that does not exist.

The outbox pattern removes the gap. Inside the same transaction as the state change, insert a row into an outbox table describing the event. A separate relay process reads unpublished rows in order and publishes them, marking each sent.

SQL
BEGIN;  UPDATE payment SET status = 'CAPTURED' WHERE id = 88123;  INSERT INTO ledger_entry (...) VALUES (...), (...);  INSERT INTO outbox (aggregate_id, type, payload)       VALUES (88123, 'PaymentCaptured', '{"amount":4999,"currency":"USD"}');COMMIT;

Now the state change and the intent to publish are atomic. The relay publishes at least once, so consumers must be idempotent — the same requirement as the message queue's delivery semantics.

Two implementation notes worth having ready. The relay can poll the table, which is simple and adds latency equal to the poll interval, or tail the database's replication log, which is lower latency and more operationally involved. And ordering matters for a single payment: publish outbox rows for one aggregate in insertion order, or a consumer can see PaymentCaptured before PaymentAuthorised.

What the customer sees during the inconsistent window

Be explicit, because interviewers ask. Between authorisation and order creation there is a window — typically under a second, occasionally minutes if something is retrying — where the customer's card shows a pending charge and the order does not exist yet.

The correct interface behaviour is to show the payment as processing and the order as pending, never as failed, and to converge. The wrong behaviour is to show an error, because the customer then retries and creates the duplicate that the idempotency key in the first half of this lesson exists to prevent. Interface design and consistency model are connected here, and saying so is a mark of experience.