System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

Estimation in practice: a worked example, rounding and sanity checks


Knowing the four quantities is not the same as producing them at interview pace. The arithmetic is not hard. Doing it in your head, out loud, while someone watches, is hard.

This lesson closes that gap in three parts. First, the whole thing done end to end on a Twitter-scale feed, at interview pace, with every assumption spoken. Second, the rounding rules that make that pace possible. Third, the most valuable estimation skill of all: recognising within five seconds that a number is absurd.

The assumptions, stated out loud

The whole estimate, at interview pace150 M post dailyx2 posts / 86,400 s3.5 K writes/sRead:write is 100:13.5 K x 100350 K reads/sA tweet is300 bytes300 M/day x 300 B90 GB/day, 165 TBEgress dominates350 K/s x 300 Babout 105 MB/s20 percent hot set90 GB x 0.218 GB in memoryAssumptionWorkingResultWrite QPSRead QPSStorageBandwidthCacheAll figures invented, one significant figure throughout.
Every row is one multiplication, and the last column is what the architecture is then argued against.

All figures in this worked example are invented for the exercise.

"I'll assume 300 million monthly active users, and that half of them are daily active — so 150 million daily active users. I'll assume the average user posts twice a day, and opens their feed ten times a day, loading twenty posts each time. Tell me if any of that is off and I'll redo it."

That is four assumptions in fifteen seconds. Everything below follows from them.

Write QPS

150M users × 2 posts = 300 million posts a day

300M ÷ 10^5 seconds = 3,000 posts per second average

Peak at 3× = 9,000 posts per second

Three thousand writes per second is already past what a single relational leader will absorb comfortably (Sharding, and what it takes away). Sharding is now justified rather than assumed.

Read QPS

150M users × 10 feed opens = 1.5 billion feed requests a day

1.5B ÷ 10^5 = 15,000 feed requests per second, peak 45,000

Read-to-write is 5:1 at the request level. But each feed request returns twenty posts, so at the level of posts delivered it is 100:1 — and that is the number that makes fan-out the central question of the design (Section 13, news feed).

Storage

Post text ~300 bytes, metadata and identifiers ~200 bytes → say 1 KB per post with indexes.

300M × 1 KB = 300 GB a day of text

× 365 = ~110 TB a year → 550 TB over five years, before replication

Media is the bigger half:

Assume 10% of posts carry an image at 200 KB.

30M × 200 KB = 6 TB a day → ~2.2 PB a year → 11 PB over five years

Two conclusions fall out immediately: text goes in a sharded store, and media goes in blob storage behind a CDN, never in the database.

Bandwidth

Ingress: 6 TB of media a day ÷ 10^5 s = 60 MB/s, or about 0.5 Gbps. Manageable.

Egress, text: 15,000 requests/s × 20 posts × 1 KB = 300 MB/s ≈ 2.4 Gbps.

Egress, media: far larger, and served entirely by the CDN.

Two gigabits per second of text alone means the read path needs a cache in front of it.

Cache

Assume users mostly read recent posts — say the last three days covers most reads.

300 GB/day × 3 = 900 GB of text. Cache only the hot 20%: ~180 GB.

At 64 GB usable per node, that is 3 nodes, or 6 with a replica each.

Six cache machines to absorb most of 45,000 peak requests per second. That is a cheap and defensible trade, and now you can say so with a number attached.

The summary line

At the end, say the whole thing in one breath — this is the sentence the interviewer writes down:

"So: 9,000 writes and 45,000 reads per second at peak, 550 TB of text and 11 PB of media over five years, about 2.4 Gbps of text egress, and a 180 GB hot set that fits in six cache nodes. The read-to-write ratio at post level is 100:1, so the design should optimise reads even if that makes writes more expensive."

Elapsed: under three minutes. Everything after this can point back at it. The conclusion "optimise reads even if writes get more expensive" is worth more than any individual number.

Rounding aggressively: the rules

That pace is only possible with aggressive rounding.

Round before you multiply, not afterThe rules• A day is 100,000 seconds, not 86,400• Powers of ten and one leading digit• Carry the exponent, not the digitsWhy it is safe• A 15 percent error changes no decision• The design is the same at 3 K or 3.5 K• Speed buys minutes for the architecture
Precision you cannot defend costs time and buys nothing, because no design choice turns on the second digit.
  1. One significant figure. 86,400 becomes 100,000. 365 becomes 400 — or skip it and use "×400 for a year" then adjust. 512 bytes becomes 500.
  2. Powers of ten only. Convert everything to scientific notation in your head: 150 million is 1.5 × 10^8. Multiplying and dividing powers of ten is adding and subtracting exponents, which is the operation you can do while talking.
  3. Never carry a decimal. If a division gives 3.7, say "about 4" and move on.
  4. Round in the direction you can defend. Rounding traffic up and capacity down leaves you with headroom, which is the safer error.

The same calculation, slowly and quickly

Slowly, the way it goes wrong:

"150 million users times 2 posts is 300 million posts. Divided by 86,400 seconds… so 300,000,000 divided by 86,400… let me see, 86,400 times 3,000 is 259 million, so it's a bit more than 3,000, maybe 3,472… let me get that more precisely…"

Ninety seconds gone, and the extra precision is worthless because the input was a guess.

Quickly:

"1.5 × 10^8 users, 2 posts each, is 3 × 10^8 posts a day. A day is 10^5 seconds. So 3 × 10^3 — three thousand writes per second. Peak, call it ten thousand."

Twelve seconds. Same conclusion. The second version also leaves the interviewer's attention on the design rather than on your mental arithmetic.

Why the precision genuinely does not matter

Your input was "assume half of monthly actives are daily active". That assumption could be off by a factor of two in either direction. Computing 3,472 from it is false precision — the error in the assumption is a hundred times the error from rounding.

What the number is for is choosing between options that differ by orders of magnitude: one machine or a hundred, one database or sixteen shards, cache or no cache. Those decisions do not change between 3,000 and 3,472. They change between 3,000 and 300,000.

The three-minute budget

Roughly: 30 seconds stating assumptions, 30 seconds on QPS, 45 on storage, 30 on bandwidth, 30 on cache, 15 on the summary. If you overrun, drop bandwidth — it is the one most often inferable from the storage figure.

If the interviewer says "you can skip the estimation", believe them and skip it. Some interviewers care about it much less than others, and spending three minutes they explicitly declined is a poor read of the room.

Sanity-checking: reference points to check against

Speed creates its own risk: a fast wrong number. The most valuable estimation skill is not producing a number. It is recognising within five seconds that a number is absurd.

Reference points to check againstIs thisnumber absurd?A day is 100,000 s1 M/day is 12 per sSSD read: 0.1 msDisk seek: 10 msCross-region: 100 ms1 Gbps is 125 MB/s
If a small product needs four hundred machines, an exponent slipped — and catching that is worth more than the arithmetic.

Keep a short list of anchors. When an estimate lands far outside them, something is wrong.

AnchorApproximate value
Simple requests one server core can handlethousands per second
Queries a single well-tuned relational database absorbsa few thousand per second
Writes a single relational leader absorbs comfortablya few thousand per second
Operations a single in-memory cache node handleson the order of 100,000 per second
Sequential throughput of one NVMe SSD1–7 GB/s
Sequential throughput of one spinning disk100–200 MB/s
1 Gbps of network125 MB/s
Memory on one large cloud instancehundreds of GB to a few TB
Peak request rate of a very large consumer serviceorder of 10^5–10^6 per second

The three questions to ask of any result

1. How many machines does this imply? Divide your load by the anchor. If the answer is "0.02 machines", the system is small and your design should say so. If it is "400,000 machines", you have made an arithmetic error — that is larger than most companies' entire fleet.

2. Does it match something I know? If your estimate for a mid-sized product produces more traffic than a global service, one of your assumptions is wrong by an order of magnitude. Usually it is actions-per-user-per-day.

3. Do the units survive? Bytes per second, requests per second, bytes per request. Most absurd results come from a unit slip — dividing by 86,400 twice, or forgetting that a megabyte is 10^6 bytes and confusing it with a megabit.

What to do when the number is absurd

Say so. Out loud, immediately:

"That gives me 400,000 servers, which cannot be right — that's larger than the whole company. Let me check… I divided by seconds twice. It's 4,000 servers, which is still large but plausible for a service of this size."

This is a good moment in an interview, not a bad one. You have demonstrated that you check your own work, which is precisely the behaviour that makes a senior engineer trustworthy. Silently carrying a wrong number into the design is the failure; catching it is the skill.

The final check: does the number change the design?

If the estimate would produce the same architecture whether it were 3,000 or 30,000, it did not need to be precise. If crossing 10,000 would flip you from one database to sixteen, that is the number to spend an extra thirty seconds on — and to say so: "I'm near the threshold where this stops fitting on one machine, so I'd want a real measurement before committing."