Course Content
System Design Interview
31 sections · 71 lessons
Payment System: reconciliation, failure handling and follow-ups
Every mechanism so far defends against a failure you anticipated. Reconciliation is what makes the system trustworthy in the presence of failures you did not — and it is the topic that free material almost never covers.
This lesson covers reconciliation first, then the failure the whole design has been building towards — a provider call whose outcome you cannot know — and finally refunds, chargebacks, PCI DSS and the questions interviewers ask to close the round.
What reconciliation is
Once a day, the provider sends a settlement file: an authoritative list of every transaction that actually settled, with amounts, fees, and identifiers. Reconciliation compares that file, line by line, against your ledger.
It is the same idea as ad click reconciliation — an independent recomputation used as a cross-check — except the independent source is not your own batch job. It is the counterparty's record, which means it fails in genuinely uncorrelated ways.
The five mismatch classes
Every difference falls into one of these, and having the taxonomy ready is what makes the answer sound practised rather than improvised.
| Class | What it means | Typical cause | Repair |
|---|---|---|---|
| In provider, not in us | They charged; we have no record | Our service died after the provider succeeded | Create the missing ledger entries; investigate whether the customer got the goods |
| In us, not in provider | We recorded a charge that never settled | We recorded optimistically, or the provider dropped it | Reverse the internal entries; alert if frequent |
| Amount mismatch | Same transaction, different value | Currency conversion, partial capture, fee accounting | Adjust to the provider's figure; the counterparty's amount is authoritative for settled money |
| Status mismatch | We say captured, they say refunded | A webhook was missed or arrived out of order | Re-query the provider for current status and apply |
| Duplicate | Two internal records for one settlement | An idempotency defence failed somewhere | Void one, and treat it as a defect to investigate, not only to repair |
Sizing the exception workload
10 million payments a day at a 0.01% mismatch rate is 1,000 exceptions per day. That number decides whether reconciliation is a job or a department.
At one minute of human attention each, 1,000 exceptions is over 16 person-hours daily — unsustainable. So the design must auto-repair the well-understood classes and escalate only the rest:
- Auto-repair amount mismatches from known fee or rounding rules, and status mismatches resolved by re-querying the provider.
- Auto-repair with notification for "in provider, not in us" where the payment identifier matches an internal record in a pending state — the common crash case.
- Escalate to a human for duplicates, for anything with no matching internal record at all, and for anything above a value threshold.
If auto-repair handles 95%, 50 exceptions a day reach a person. That is a workable operations process, and stating the arithmetic is how you show the design is operable rather than merely correct.
Corrections are new entries, never edits
A ledger entry, once written, is never modified or deleted. A correction is a new entry that reverses or adjusts the original and references it.
This is not bureaucratic caution. If entries can be edited, then the ledger's history depends on who edited what and when, and no statement about a past date can be trusted. If entries are immutable, the ledger at any past instant is exactly reconstructible, which is what an auditor, a regulator, and a bug investigation all require.
Reconcile against more than one source
The settlement file is one counterparty. A mature system reconciles several ways: internal ledger against the provider, provider against the bank account the money actually landed in, and the sum of merchant payable balances against the platform's own bank balance.
Each comparison catches a different class of error. The bank comparison in particular catches things no provider file can — money that settled to the wrong account, or fees deducted outside the file.
The timeout where the outcome is genuinely unknown
Reconciliation is the safety net. This is the failure it most often catches, and the most important failure in the system. Your service calls the provider to authorise; the connection times out after 30 seconds. Three possibilities, indistinguishable from where you stand: the request never arrived, it arrived and was declined, or it arrived and succeeded and the response was lost.
Retrying blindly is wrong. If it succeeded, you charge twice. Failing the payment is wrong. If it succeeded, the customer is charged for an order marked failed.
The correct handling is a distinct state and a defined procedure:
- Move the payment to
IN_DOUBTand record the timeout with its timestamp. Never leave it inPENDING, because pending is indistinguishable from "not started". - Query the provider by your own reference — every serious provider supports lookup by the merchant's reference identifier for exactly this situation. This is why step 2 of The payment flow sends your payment identifier with the request.
- If the query says approved, transition to
AUTHORISEDand continue. If declined or not found, transition toDECLINED. - If the query itself fails, retry it with backoff for a bounded period, then reverse: send an explicit void or refund for your reference, which is a no-op if nothing was charged and a correction if something was.
- If everything fails, leave it
IN_DOUBTand let reconciliation resolve it the next day. It will appear as an "in provider, not in us" or "in us, not in provider" mismatch and be repaired.
Notice the shape: never guess, query first, reverse second, reconcile last. Give that sequence in the interview and the follow-up is usually finished.
Refunds and partial refunds
A refund is a new payment in the opposite direction, not an undo. It gets its own identifier, its own idempotency key, its own ledger entries, and its own state machine.
Partial refunds add an invariant: the sum of refunds against a payment must never exceed the captured amount. Enforce it in the database — a check against a running total, or a constraint that the sum of refund entries plus the capture entry cannot go negative — rather than only in application code, for the same reason as the hotel booking lesson: a constraint survives a new code path, an application check does not.
Multi-currency refunds carry a subtlety worth naming. A payment captured in one currency and refunded weeks later at a different exchange rate leaves a difference. Somebody absorbs it, the policy is a business decision, and the ledger must record the difference explicitly as a foreign-exchange gain or loss entry rather than letting it vanish into a rounding discrepancy.
Chargebacks
A cardholder disputes a charge with their issuer. The issuer pulls the funds back from the merchant, often months later, and the merchant may contest with evidence within a deadline.
Architecturally this means: a long-lived state machine per dispute, deadline tracking with alerting, storage of evidence artefacts, ledger entries for the reversal and for the dispute fee, and the understanding that the outcome can arrive long after everyone has forgotten the order. It is a workflow system attached to the payment system, and describing it as one is the right answer.
PCI DSS, and the parts that change the architecture
The Payment Card Industry Data Security Standard governs handling of card data. Its practical architectural consequence is one sentence: the cheapest way to comply is to never touch a card number.
Two mechanisms achieve that:
- Hosted fields or a redirect. The card number is entered into an iframe or page served by the provider, so it never reaches your servers or your browser JavaScript.
- Tokenisation. The provider returns an opaque token representing the card, which you store and use for future charges. The token is useless to an attacker outside your provider account.
The consequence for the design is that your database stores a token, an expiry, a brand, and the last four digits — enough for a user interface and support — and never a full card number. That shrinks the systems in regulatory scope from "everything" to "the checkout page", which is the difference between a manageable audit and an enormous one.
Compliance requirements evolve, so treat specifics as something to verify rather than to quote.