Course Content
System Design Interview
31 sections · 71 lessons
Estimation: the numbers to know and the four quantities to compute
Back-of-the-envelope estimation is arithmetic done in your head, to one significant figure, to turn a vague prompt into concrete constraints. It takes three minutes and it changes how every later decision reads.
This lesson gives you everything you need to produce the numbers. First, why estimation earns points at all. Second, the small set of figures to know without thinking — everything else in estimation is multiplication. Third, the four quantities you always compute, in the order that lets each one feed the next.
What estimation does for the interview
A prompt like "design a photo-sharing service" contains no numbers, so every architectural choice you make afterwards is unjustified. Estimation manufactures the justification.
Consider two candidates answering "should this data live on one database?"
Candidate A: "One database will not scale, so I would shard it."
Candidate B: "We're storing 300 million rows a day at about 1 KB each — that's 300 GB a day, 110 TB a year. A single machine's disk is gone in about a week, so this has to be sharded, and I'll shard on user ID because the read pattern is per-user."
Same conclusion. The first is a reflex; the second is a derivation. Only the second tells the interviewer that the candidate would make the opposite call if the numbers were smaller, which is the thing being assessed (How you are scored, and how level is decided).
It also stops you over-designing
Estimation protects you in both directions. Run the numbers on a corporate expenses tool used by 5,000 employees and you get about 50,000 requests a day — under one per second. A candidate who proposes Kafka, sixteen shards, and a service mesh for that has demonstrated worse judgement than one who says "this runs on two application servers and one database, and here is what I would monitor to know when that stops being true."
Interviewers notice both failures. The over-engineering failure is more common and, at senior level, more damaging.
The constraints it produces
Four numbers, computed in a fixed order (the second half of this lesson walks through each), each of which forces a design decision:
| Number | What it decides |
|---|---|
| Queries per second, average and peak | How many servers, and whether one database can absorb the writes |
| Storage per day and over five years | One database or many; hot storage or tiered |
| Bandwidth in and out | Whether a CDN is required rather than optional |
| Memory for the cache | How many cache nodes, and whether the hot set fits at all |
The two failure modes
Skipping it. The design proceeds on vibes and every choice is unbacked. The interviewer starts asking "why?" and there is nothing underneath.
Drowning in it. Ten minutes computing the exact size of a timestamp column. You have spent a fifth of the interview and produced a number nobody needed. The lesson on rounding is about staying fast.
It is a spoken exercise, not a silent one
State every assumption out loud as you make it: "I'll assume 100 million daily active users and that 10% of them post once a day — stop me if that's wrong." Two things happen. The interviewer often corrects you, which is free information and exactly the collaboration they want to see. And if they do not, the assumption is now shared, so your later numbers cannot be called wrong — only your arithmetic can.
The numbers to memorise: powers of two
There is a small set of figures you should know without thinking. The first group is powers of two, which is really data-size vocabulary.
| Power | Approximate value | Name | Common use |
|---|---|---|---|
| 2^10 | 1 thousand | 1 KB (kilobyte) | A short text record |
| 2^20 | 1 million | 1 MB (megabyte) | A photo thumbnail, a page of logs |
| 2^30 | 1 billion | 1 GB (gigabyte) | Memory on a small server |
| 2^40 | 1 trillion | 1 TB (terabyte) | One machine's disk |
| 2^50 | 1 quadrillion | 1 PB (petabyte) | A large company's data lake |
Treating 2^10 as exactly one thousand introduces a 2.4% error and saves you from doing long division under pressure. Take the error.
Also worth knowing: 32 bits addresses about 4 billion values, and 64 bits about 1.8 × 10^19. That single fact answers most "how many bits do we need?" questions (Section 9).
Time
- 86,400 seconds in a day. Round it to 100,000 — that is 10^5, and it makes every division a shift of the decimal point.
- 1 million requests a day ≈ 10 per second (12 with the exact figure).
- 1 billion requests a day ≈ 10,000 per second.
- 2.5 million seconds in a month; 31.5 million in a year (call it 3 × 10^7).
Latency, in orders of magnitude
| Operation | Order of magnitude |
|---|---|
| Main memory reference | ~100 nanoseconds (ns) |
| Compress 1 KB with a fast algorithm | ~1–2 microseconds (µs) |
| Send 1 KB over a 1 Gbps network | ~10 µs |
| Read 4 KB randomly from an SSD | ~100 µs |
| Round trip within the same datacentre | ~0.5 milliseconds (ms) |
| Read 1 MB sequentially from an SSD | ~1 ms |
| Disk seek on a spinning disk | ~10 ms |
| Round trip across a continent or ocean | ~100–150 ms |
Two derived facts carry most of the weight: a network round trip inside a datacentre costs about as much as reading a megabyte from disk, and a cross-continent round trip is roughly 300× a local one, which is why replication across regions changes a design and replication within a region does not.
Availability
| Target | Downtime per year | Per month |
|---|---|---|
| 99% | 3.65 days | 7.2 hours |
| 99.9% ("three nines") | 8.8 hours | 43 minutes |
| 99.99% | 53 minutes | 4.3 minutes |
| 99.999% | 5.3 minutes | 26 seconds |
Useful because a candidate who says "highly available" and one who says "99.9%, so we can absorb a 40-minute monthly outage but not a daily one" are not saying the same thing.
The four quantities you always compute
With the reference card in your head, the computation itself is always the same four quantities, always in this order, because each one feeds the next.
1. Queries per second (QPS), average and peak
Start from users, not from requests.
daily active users × actions per user per day ÷ 100,000 seconds = average QPS
Then apply a peak multiplier. Traffic is not flat: a consumer app in one country might see three to five times its average in the evening; a global app is flatter, maybe two to three times. State which you are assuming.
peak QPS = average QPS × 3 (state the multiplier out loud)
Compute writes and reads separately. The ratio between them is often the single most important number in the design. A 100:1 read-to-write ratio (Section 10, URL shortener) says "cache aggressively and add read replicas". A 1:1 ratio says "the write path is your problem".
2. Storage, per day and over five years
new objects per day × bytes per object = storage per day
Then multiply by 365 and by the retention period. Five years is the conventional horizon — long enough to expose a problem, short enough to stay believable.
Two adjustments people forget:
- Replication. Storing three copies triples the bill. If you say "replication factor 3", multiply.
- Indexes and overhead. A 200-byte row is not 200 bytes on disk. Multiplying raw size by 1.5 to 2 is a reasonable, statable allowance.
3. Bandwidth, in and out
ingress = writes per second × bytes per write
egress = reads per second × bytes per read
Egress is usually the dramatic one, because the read multiplier applies to it. For any system serving images or video, this number is what makes a CDN mandatory rather than an optimisation (Section 16, Design YouTube, is built around it).
Convert to a unit you can sanity-check: 1 Gbps = 125 MB/s. If your egress figure is larger than a few gigabits per second, you are describing a CDN-shaped system.
4. Memory for the cache
cache size = the portion of data that is hot × bytes per item
The usual working assumption is a heavy-tailed access distribution — often stated as 80% of requests hitting 20% of the data. Say that you are assuming it rather than presenting it as fact; it is a rule of thumb, not a measurement of your system.
Then check it against real hardware. A cache node might have 64 GB usable. A 300 GB hot set is five nodes, or ten with replication. A 30 TB hot set does not fit in a cache and you have learnt something important about the design.
The worked shape
For a system with 10 million daily active users writing twice a day and reading fifty times a day, at 1 KB per item:
- Writes: 20M/day ÷ 10^5 = 200 writes/s, peak 600
- Reads: 500M/day ÷ 10^5 = 5,000 reads/s, peak 15,000
- Storage: 20M × 1 KB = 20 GB/day → 7.3 TB/year → ~37 TB over five years, ×3 for replication ≈ 110 TB
- Egress: 5,000/s × 1 KB = 5 MB/s (unremarkable — no CDN needed for this payload)
- Cache: 20% of a week's data = 20 GB/day × 7 × 0.2 = 28 GB — one node, comfortably
Four lines of arithmetic, and the design is already constrained.