System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

Estimation: the numbers to know and the four quantities to compute


Back-of-the-envelope estimation is arithmetic done in your head, to one significant figure, to turn a vague prompt into concrete constraints. It takes three minutes and it changes how every later decision reads.

This lesson gives you everything you need to produce the numbers. First, why estimation earns points at all. Second, the small set of figures to know without thinking — everything else in estimation is multiplication. Third, the four quantities you always compute, in the order that lets each one feed the next.

What estimation does for the interview

The two ways estimation goes wrongSkipping the numbers• Choices sound like personal preference• No argument for the cache size• Over-designs for traffic that is absentDrowning in the numbers• Three minutes becomes twelve• Two significant figures nobody needs• The architecture never gets drawn
The number is not the deliverable; the constraint it places on the next decision is.

A prompt like "design a photo-sharing service" contains no numbers, so every architectural choice you make afterwards is unjustified. Estimation manufactures the justification.

Consider two candidates answering "should this data live on one database?"

Candidate A: "One database will not scale, so I would shard it."

Candidate B: "We're storing 300 million rows a day at about 1 KB each — that's 300 GB a day, 110 TB a year. A single machine's disk is gone in about a week, so this has to be sharded, and I'll shard on user ID because the read pattern is per-user."

Same conclusion. The first is a reflex; the second is a derivation. Only the second tells the interviewer that the candidate would make the opposite call if the numbers were smaller, which is the thing being assessed (How you are scored, and how level is decided).

It also stops you over-designing

Estimation protects you in both directions. Run the numbers on a corporate expenses tool used by 5,000 employees and you get about 50,000 requests a day — under one per second. A candidate who proposes Kafka, sixteen shards, and a service mesh for that has demonstrated worse judgement than one who says "this runs on two application servers and one database, and here is what I would monitor to know when that stops being true."

Interviewers notice both failures. The over-engineering failure is more common and, at senior level, more damaging.

The constraints it produces

Four numbers, computed in a fixed order (the second half of this lesson walks through each), each of which forces a design decision:

NumberWhat it decides
Queries per second, average and peakHow many servers, and whether one database can absorb the writes
Storage per day and over five yearsOne database or many; hot storage or tiered
Bandwidth in and outWhether a CDN is required rather than optional
Memory for the cacheHow many cache nodes, and whether the hot set fits at all

The two failure modes

Skipping it. The design proceeds on vibes and every choice is unbacked. The interviewer starts asking "why?" and there is nothing underneath.

Drowning in it. Ten minutes computing the exact size of a timestamp column. You have spent a fifth of the interview and produced a number nobody needed. The lesson on rounding is about staying fast.

It is a spoken exercise, not a silent one

State every assumption out loud as you make it: "I'll assume 100 million daily active users and that 10% of them post once a day — stop me if that's wrong." Two things happen. The interviewer often corrects you, which is free information and exactly the collaboration they want to see. And if they do not, the assumption is now shared, so your later numbers cannot be called wrong — only your arithmetic can.

The numbers to memorise: powers of two

There is a small set of figures you should know without thinking. The first group is powers of two, which is really data-size vocabulary.

PowerApproximate valueNameCommon use
2^101 thousand1 KB (kilobyte)A short text record
2^201 million1 MB (megabyte)A photo thumbnail, a page of logs
2^301 billion1 GB (gigabyte)Memory on a small server
2^401 trillion1 TB (terabyte)One machine's disk
2^501 quadrillion1 PB (petabyte)A large company's data lake

Treating 2^10 as exactly one thousand introduces a 2.4% error and saves you from doing long division under pressure. Take the error.

Also worth knowing: 32 bits addresses about 4 billion values, and 64 bits about 1.8 × 10^19. That single fact answers most "how many bits do we need?" questions (Section 9).

Time

  • 86,400 seconds in a day. Round it to 100,000 — that is 10^5, and it makes every division a shift of the decimal point.
  • 1 million requests a day ≈ 10 per second (12 with the exact figure).
  • 1 billion requests a day ≈ 10,000 per second.
  • 2.5 million seconds in a month; 31.5 million in a year (call it 3 × 10^7).

Latency, in orders of magnitude

OperationOrder of magnitude
Main memory reference~100 nanoseconds (ns)
Compress 1 KB with a fast algorithm~1–2 microseconds (µs)
Send 1 KB over a 1 Gbps network~10 µs
Read 4 KB randomly from an SSD~100 µs
Round trip within the same datacentre~0.5 milliseconds (ms)
Read 1 MB sequentially from an SSD~1 ms
Disk seek on a spinning disk~10 ms
Round trip across a continent or ocean~100–150 ms

Two derived facts carry most of the weight: a network round trip inside a datacentre costs about as much as reading a megabyte from disk, and a cross-continent round trip is roughly 300× a local one, which is why replication across regions changes a design and replication within a region does not.

Availability

TargetDowntime per yearPer month
99%3.65 days7.2 hours
99.9% ("three nines")8.8 hours43 minutes
99.99%53 minutes4.3 minutes
99.999%5.3 minutes26 seconds

Useful because a candidate who says "highly available" and one who says "99.9%, so we can absorb a 40-minute monthly outage but not a daily one" are not saying the same thing.

Numbers worth memorising before the interviewLatencyL1 cache reference1 nsMain memory reference100 nsSSD random read100 µsRound trip within a datacentre500 µsDisk seek10 msPacket India → US → India~150 mseach row is roughly 100× the one abovePowers of two2¹⁰1 thousand — KB2²⁰1 million — MB2³⁰1 billion — GB2⁴⁰1 trillion — TB2⁵⁰1 quadrillion — PBclose enough for an interview; nobody wants 1,048,576TimeSeconds in a day86,400 ≈ 10⁵Seconds in a month2.6 × 10⁶Seconds in a year3.2 × 10⁷1M/day≈ 12 per second1B/day≈ 12,000 per seconddivide daily volume by 10⁵ for a rough QPSthe estimation sequencedaily active users× actions per user= writes per day÷ 86,400 = QPS× peak factor (2–10×)× bytes per item = storageState the assumption out loud, round hard, and keep one significant figure. An interviewer is checking whether the number changes your design, not whether it isright to three decimal places.Peak is not average: a 5× peak factor is the difference between a design that survives launch day and one that does not.
Six latency rows, five storage rows and one sequence — enough to size almost any system on a whiteboard.

The four quantities you always compute

With the reference card in your head, the computation itself is always the same four quantities, always in this order, because each one feeds the next.

Each number feeds the nextQPS,average and peakStorage: dayand 5 yearsBandwidthin and outCache:the hot setPeak is roughly twice average unless the product says otherwise.
The order is not arbitrary — bandwidth needs the QPS, and the cache size needs the daily storage.

1. Queries per second (QPS), average and peak

Start from users, not from requests.

daily active users × actions per user per day ÷ 100,000 seconds = average QPS

Then apply a peak multiplier. Traffic is not flat: a consumer app in one country might see three to five times its average in the evening; a global app is flatter, maybe two to three times. State which you are assuming.

peak QPS = average QPS × 3 (state the multiplier out loud)

Compute writes and reads separately. The ratio between them is often the single most important number in the design. A 100:1 read-to-write ratio (Section 10, URL shortener) says "cache aggressively and add read replicas". A 1:1 ratio says "the write path is your problem".

2. Storage, per day and over five years

new objects per day × bytes per object = storage per day

Then multiply by 365 and by the retention period. Five years is the conventional horizon — long enough to expose a problem, short enough to stay believable.

Two adjustments people forget:

  • Replication. Storing three copies triples the bill. If you say "replication factor 3", multiply.
  • Indexes and overhead. A 200-byte row is not 200 bytes on disk. Multiplying raw size by 1.5 to 2 is a reasonable, statable allowance.

3. Bandwidth, in and out

ingress = writes per second × bytes per write

egress = reads per second × bytes per read

Egress is usually the dramatic one, because the read multiplier applies to it. For any system serving images or video, this number is what makes a CDN mandatory rather than an optimisation (Section 16, Design YouTube, is built around it).

Convert to a unit you can sanity-check: 1 Gbps = 125 MB/s. If your egress figure is larger than a few gigabits per second, you are describing a CDN-shaped system.

4. Memory for the cache

cache size = the portion of data that is hot × bytes per item

The usual working assumption is a heavy-tailed access distribution — often stated as 80% of requests hitting 20% of the data. Say that you are assuming it rather than presenting it as fact; it is a rule of thumb, not a measurement of your system.

Then check it against real hardware. A cache node might have 64 GB usable. A 300 GB hot set is five nodes, or ten with replication. A 30 TB hot set does not fit in a cache and you have learnt something important about the design.

The worked shape

For a system with 10 million daily active users writing twice a day and reading fifty times a day, at 1 KB per item:

  • Writes: 20M/day ÷ 10^5 = 200 writes/s, peak 600
  • Reads: 500M/day ÷ 10^5 = 5,000 reads/s, peak 15,000
  • Storage: 20M × 1 KB = 20 GB/day → 7.3 TB/year → ~37 TB over five years, ×3 for replication ≈ 110 TB
  • Egress: 5,000/s × 1 KB = 5 MB/s (unremarkable — no CDN needed for this payload)
  • Cache: 20% of a week's data = 20 GB/day × 7 × 0.2 = 28 GB — one node, comfortably

Four lines of arithmetic, and the design is already constrained.