Course Content
System Design Interview
31 sections · 71 lessons
Payment System: requirements, scale and the payment flow
"Design a payment system." Part V of this course (Sections 28–30) opens here, and it has a different character from everything before it. One sentence carries the change:
Availability is negotiable. Correctness is not.
This lesson sets up the problem — the priorities, the vocabulary, the numbers — and builds the payment flow. The second lesson goes deep on double-payment prevention and consistency across services; the third covers reconciliation, failure handling and the follow-up questions.
Why Part V feels different
In Section 13, a feed that is briefly stale is fine. In Section 22, a dropped metric is invisible. In Section 26, an object is durable to eleven nines and that is considered excellent.
None of that transfers. If a payment system is unavailable for ten minutes, customers are annoyed and the money is still right. If a payment system is available and wrong for ten minutes, you have double-charged people, lost money you cannot account for, and created a regulatory problem. The failure modes are not symmetric, so the design priorities invert: when in doubt, stop rather than guess.
Three concrete consequences you will see throughout Sections 28 to 30:
- Refusing to serve is an acceptable answer. A payment that cannot be processed safely is declined, and declining is a normal, well-handled outcome.
- Every state change is recorded, not overwritten. An audit trail is a functional requirement, not observability.
- Nothing is trusted until it is reconciled against an independent record. Reconciliation is the most important part of this section for exactly this reason.
Vocabulary you need before the first question
The questions that shape everything after
- Which payment methods? Cards, bank transfers, wallets, and buy-now-pay-later each have different flows, different timings, and different failure modes. Cards are the hardest and the usual choice.
- Are we the merchant, or are we building the processor? Building a payment page for one business and building a payment platform for thousands are different systems. Assume the first unless told otherwise.
- Do we hold funds, or does money move directly? Holding funds makes us a wallet, which is Section 29 (Design a Digital Wallet), and brings licensing obligations.
- Multi-currency? If yes, every amount needs a currency, foreign-exchange rates need a quote-and-hold mechanism, and rounding rules become a correctness issue.
- Refunds and chargebacks in scope? Chargebacks in particular have a long, stateful, deadline-driven workflow that is easy to underestimate.
- What is the regulatory scope? For cards this means PCI DSS — the Payment Card Industry Data Security Standard — and its requirements change the architecture, as Failure handling and follow-ups shows.
The assumptions this section uses
| Question | Assumption |
|---|---|
| Methods | Cards, via one or more external payment service providers |
| Role | We are the merchant's platform, not the card network |
| Funds | We do not hold customer balances; money flows to the merchant |
| Currency | Multi-currency, with amounts stored in minor units |
| Refunds | Full and partial refunds, plus chargeback handling |
| Regulation | PCI DSS scope minimised by never touching card numbers |
Requirements and scale
Functional requirements
- Accept a payment for an order, using a stored or newly entered card.
- Authorise, then capture, then record the result against the order.
- Support full and partial refunds.
- Maintain a ledger of every money movement.
- Handle asynchronous notifications from the payment service provider.
- Reconcile internal records against the provider's settlement file daily.
- Expose payment status and history to customer support and finance.
Non-functional requirements
These are the requirements that dominate, and they should be stated as absolutes:
- No double charge. A customer is charged once per intent, regardless of retries, network failures, or duplicate clicks.
- No lost payment. A payment that succeeded at the provider is recorded internally, always.
- A complete audit trail. Every state transition is recorded immutably with who, what, when, and why.
- Consistency over availability. Under partition or failure, decline rather than proceed.
Only then the conventional ones: authorisation round trip within a few seconds, since a user is waiting; and the ability to survive a provider outage by failing over.
The arithmetic
Assume a large e-commerce platform: 10 million payments per day. Invented, but the right order of magnitude.
Rate. 10,000,000 ÷ 100,000 s = 100 payments per second average. Retail traffic peaks hard — a sale event or a Friday evening can be 10×, so design for 1,000 per second.
Ledger volume. Each payment produces several ledger entries: the authorisation, the capture, the fee, the payable to the merchant. Call it 6 entries per payment on average including refunds and adjustments. 10 million × 6 = 60 million ledger entries per day.
Storage. At 300 bytes per entry, that is 18 GB per day, 6.6 TB per year. Financial records are commonly retained for seven years or more depending on jurisdiction — treat the exact period as a legal question rather than an engineering one — so plan for 45 to 65 TB. Modest, and it must never be deleted casually.
Provider calls. Each payment is at least one outbound authorisation and one capture, so 2,000 external calls per second at peak, each taking a few hundred milliseconds. At 300 ms per call, sustaining 2,000 per second requires 600 concurrent in-flight requests. That number matters: it sizes the connection pool, and it is where a provider slowdown turns into a queue in your own system.
What the numbers rule in
- A relational database with real transactions is correct here. At 1,000 payments per second it is nowhere near its limits, and the guarantees it provides are exactly the ones required. Reaching for an eventually-consistent store in this problem is a mistake.
- Sharding is not needed yet. 60 million rows a day is large but ordinary. Partition by time for manageability, and shard by merchant or account only when growth demands it.
- Caching is mostly wrong. Cached balances and cached payment statuses are how stale money data reaches a decision. Cache reference data — merchant configuration, currency codes — and read money from the source of truth.
The payment flow
Now build the flow, starting with the version that looks right and is not.
The naive flow, and its three failure points
- The customer clicks Pay.
- The service calls the payment provider and waits.
- On success, it marks the order paid and returns.
Three specific failures break it.
The provider call times out. The service does not know whether the charge happened. Retrying may double-charge; not retrying may lose a payment the customer was charged for.
The service dies after the provider succeeds. The money moved and nothing internal records it. The customer is charged, the order is unpaid, and only a human comparing records will ever find it.
The customer double-clicks. Two requests, two charges.
Each failure points at a mechanism: idempotency keys, an outbox with asynchronous confirmation, and a ledger written in the same transaction as the state change.
The components
Payment service. Owns the payment lifecycle. Accepts a payment intent, enforces idempotency, drives state transitions, and calls the provider. It is the only component that talks to the provider.
Ledger. An append-only, double-entry record of every money movement. Never updated in place. The digital wallet section develops double-entry properly; here it is enough to know that every movement is recorded as balanced debit and credit entries, so an inconsistent state cannot be represented.
Wallet or account service. Tracks balances per party — the merchant's payable balance, the platform's fee revenue. Derived from the ledger, never independently authoritative.
Payment service provider. External. Talks to acquirers, networks, and issuers. Returns synchronous responses and sends asynchronous webhooks.
Reconciliation service. Reads the provider's daily settlement file and compares it against the ledger. See Reconciliation.
The happy path, hand-off by hand-off
- Checkout → Payment service.
POST /paymentscarrying order ID, amount, currency, a payment method token, and an idempotency key generated by the client. - Payment service → its own database. In one transaction: insert the payment record in state
PENDING, insert the idempotency key, and insert an outbox row. Commit. The payment now exists internally before any external call — so an internal record is never missing. - Payment service → provider. Authorise. The provider replies approved with an authorisation identifier, declined with a reason, or times out.
- Payment service → its own database. Transition to
AUTHORISED, write the ledger entries for the authorisation, and write an outbox event. One transaction. - Payment service → provider. Capture, either immediately or when the goods ship.
- Payment service → database. Transition to
CAPTURED, write ledger entries, emitPaymentCaptured. - Outbox relay → event log. Publishes the events. Order service marks the order paid; notification service emails a receipt; analytics counts it. None of these are in the critical path.
- Provider webhook → Payment service. Asynchronous confirmations and later events — settlement, dispute, refund completion — arrive and are applied idempotently.
- Next day: settlement file → Reconciliation. The provider's authoritative record of what actually settled is compared against the ledger.