System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

The four-step framework: deep dive, wrap-up and communication


By minute 25 you have a written contract, a diagram, an API and a data model. The second half of the hour is where level is decided.

This lesson covers it in three parts. Step 3, the deep dive: fifteen minutes on one or two components, taken far enough down that the interviewer learns what you actually know. Step 4, the wrap-up: five minutes, reliably skipped, and worth more per minute than any other part of the hour. And the thing that runs under all four steps — communication — because the design is assessed through what you say.

Step 3: who chooses the deep dive

Fifteen minutes on one componentOffer twocandidatesLet them chooseGo threelevels downName thetrade-off takenThree levels means mechanism, then failure mode, then the numbers.
This is the step that decides your level, so steer it toward the component you are genuinely deepest in.

Usually the interviewer, and their choice tells you what they want to hear. If they do not choose, propose:

"The interesting part here is the fan-out, because that's where the read-to-write ratio bites. I could also go into the storage engine choice or how we handle a hot key — which would you prefer?"

Offering a menu is better than picking silently. It signals that you can see several hard parts, and it lets the interviewer steer towards their rubric.

What "three levels deep" means

Take the cache from the caching lesson in Section 2 as the component.

  • Level 1 — what it is. "A Redis cache in front of the database using cache-aside."
  • Level 2 — the mechanism and the numbers. "Read checks the cache, miss goes to the database and populates with a 60-second TTL. At a 95% hit ratio the database sees 5% of 1,700 requests per second, so 85 queries per second instead of 1,700."
  • Level 3 — the failure mode and the repair. "The problem is a hot key expiring: 5,000 concurrent requests all miss at once and hit a database sized for 85. I'd coalesce requests so one fetches and the rest wait, jitter the TTLs so keys don't expire in lockstep, and serve the stale value while a background refresh runs."

Level 3 is where the marks are. Levels 1 and 2 are recall; level 3 is the thing that is hard to fake.

Structure to avoid rambling

Each deep dive follows the same shape and takes six to eight minutes:

  1. Name the problem in one sentence. "The hard part is that one celebrity post fans out to ten million feeds."
  2. Give two or three options with what each costs.
  3. Recommend one, with the reason and the condition that would flip it.
  4. Name the failure mode of your recommendation before the interviewer does.
  5. Attach a number to at least one claim.

If you find yourself three minutes in without having named an alternative, you are narrating rather than reasoning. Stop and name one.

Saying "I don't know" well

You will be asked something you have not thought about. There is a good version and a bad version.

Bad: inventing mechanism. Interviewers detect this quickly, because the invented answer does not connect to anything else you said, and it damages every correct thing you said earlier.

Good: bound the gap and reason from what you do know.

"I haven't implemented leader election myself, so I don't know the details of the protocol. What I know is that it needs a majority to avoid split brain, that the standard answer is a consensus algorithm like Raft, and that I'd use an existing coordination service rather than writing one. If we needed to go deeper I'd want to look it up."

That answer scores well. It is honest, it demonstrates the surrounding knowledge, and it shows the judgement not to hand-roll consensus — which is the practically correct instinct.

Step 4: why the wrap-up pays so well

Five minutes, built from the recap upOne-sentence recapBottlenecks, namedHow youwould monitor itWhat youwould do nexttopbottomRead bottom to top: recap first, roadmap last.
It is the most reliably skipped part of the hour and the highest-scoring minute-for-minute.

The interviewer has spent forty minutes watching you build something. What they still do not know is whether you can see it clearly. A candidate who critiques their own design has demonstrated the thing that actually predicts on-the-job performance: they will notice the problem before it becomes an incident.

Candidates who defend their design against every question read as attached to it. Candidates who say "yes, that's a real weakness, here's when it bites" read as senior.

What to cover, in order

1. The bottleneck you would hit first. Name one component and the load at which it breaks.

"The first thing to break is the cache tier. At 45,000 reads per second with a 95% hit ratio, one node is at the edge of its throughput, so I'd shard the cache by key and watch the per-key rate for hot spots."

2. What happens when each major component dies. Go around the diagram. The cache dies: requests fall through to the database, which is not sized for it, so we need a fill-rate limit. The queue backs up: consumers lag, and the retention window is the buffer. A database shard is unreachable: those users see errors while the rest of the system is fine, which is a better outcome than a total outage and is worth saying out loud.

3. Where the design is inconsistent, and whether users notice. "Feeds are eventually consistent, so a new post can take a few seconds to appear for followers. That's acceptable here. It would not be acceptable for the balance in a digital wallet (Section 29)."

4. What you would do with more time. Two or three items, specific: "monitoring on replication lag, a load test to find the real cache throughput, and a plan for resharding before we hit 16 shards."

The one-sentence recap

Finish by restating the design in a single sentence linking it back to the requirements from step 1. The interviewer is about to write notes; give them the sentence to write.

"So: a stateless API tier behind a load balancer, writes fanned out through a queue to a sharded store partitioned by user, reads served from a cache with a 95% hit ratio, media on a CDN — which meets the 200 ms read target and the 9,000-writes-per-second peak, with eventual consistency on the feed as the accepted trade."

Communication under uncertainty: think out loud, including the dead ends

All four steps are assessed through what you say. A strong architecture communicated badly scores lower than a modest one communicated well, and that is not unfair — it mirrors the job.

How the same design gets two scoresWhat earns marks• Saying the dead end and why you left it• Treating a hint as new information• Disagreeing with a reason, then movingWhat loses them• Long silences while you think• Defending a design after it is broken• Agreeing instantly with every suggestion
The design is assessed through what you say, which mirrors the job rather than distorting it.

Silence is unscoreable. If you need twenty seconds, narrate the twenty seconds:

"I'm weighing whether to put a queue between the API and the writes. It helps with bursts but it makes the write path asynchronous, so the client can't get a confirmed identifier back. Let me check what the requirements said about that…"

That is a better minute than a silent minute followed by the same conclusion, because the interviewer can see the reasoning and can correct a wrong premise before you build on it.

Taking a suggestion

When an interviewer says "have you considered X?", they are almost always either steering you towards their rubric or telling you something is wrong. Treat it as information.

"No, I hadn't — let me think about it. X would remove the coordination problem I described, at the cost of an extra hop on every read. Given the read-to-write ratio is 100:1, that hop is expensive, so I'd want to know how often the coordination actually fires before choosing. If it's rare, I'd keep my version; if it's common, X is better."

You have taken the suggestion seriously, evaluated it against a number, and stated the condition. That answer scores whether or not you adopt the suggestion.

Disagreeing

You are allowed to disagree, and doing it well is a strong signal — as long as it is with a reason and a stated willingness to be wrong.

"I'd push back slightly. Strong consistency here would mean a cross-region round trip on every read, about 150 ms, and the requirement was 200 ms at the 99th percentile — that leaves nothing for anything else. I'd rather stay eventually consistent and handle the stale-read case in the client. But if correctness of that specific field matters more than latency, I'd change my mind."

Never disagree by repeating yourself louder, and never dismiss a suggestion without evaluating it.

Recovering when your design is wrong

It happens. At minute 30 you realise your partitioning choice makes the main query span every shard. The instinct is to hide it. The instinct is wrong.

"I've found a problem with what I drew. I partitioned by post ID, but the dominant read is 'all posts by this user', which now hits every shard. I should partition by user ID instead. That costs me even distribution if one user posts far more than others, which I'd handle with a secondary split for those accounts. Let me redraw that part."

Six sentences. You found your own bug, explained the cause, proposed a fix, and named the fix's cost. That sequence often scores higher than never having made the error, because almost nobody demonstrates it. It is the same instinct as the wrap-up: self-critique outscores defence.