Course Content
System Design Interview
31 sections · 71 lessons
Practising alone and where to go next
Reading produces recognition. Interviews test production under time pressure with someone watching. The gap between the two is closed by practice, and most of that practice can be done alone.
This lesson covers how to practise without an interviewer, how to get real value from a mock interview, and where to go after this course — the specialisation courses, the books and papers worth reading, and the one habit that matters most.
Design the systems you already use
The best practice prompts are products you use daily, because you already know the requirements and can check your design against observable behaviour.
Try: a ride-hailing dispatch system, a food-delivery order tracker, a calendar with shared availability, a password manager's sync, a code-review platform, a ticketing system for a concert on-sale, a flight search, a podcast host, a document collaboration tool, a parcel-tracking system.
For each, run the four steps from Section 4 in the real time budget, out loud, with a timer. Out loud matters — the failure mode being trained is not thinking badly, it is thinking silently and then presenting a conclusion nobody watched you reach.
The 45-minute solo protocol
- Minutes 0–10. Write the prompt at the top of a blank page. List clarifying questions and answer them yourself, plausibly. Write functional and non-functional requirements. Compute the four estimation quantities.
- Minutes 10–25. Draw the boxes. Define the API. Sketch the data model. Keep narrating.
- Minutes 25–40. Pick the hardest component and go three levels down. If you cannot find a hard part, the requirements were too easy — add a constraint and continue.
- Minutes 40–45. Write down your own bottlenecks, failure modes, and what you would do with more time.
Then, and only then, look things up. The learning happens in the gap between what you produced and what you find, and looking up early closes that gap before it can teach you anything.
Record yourself
Uncomfortable and unusually effective. Record audio while you design, then listen back a day later with a pen. Listen for:
- Silences longer than about ten seconds. In a real interview, that is where the interviewer loses the thread and starts guessing at what you are doing.
- Decisions with no stated reason. "I'll use a message queue here" with no alternative named and no trade stated. This is the single most common weakness in mid-level answers.
- Numbers you asserted instead of computing. "That won't scale" is not a claim; it is a mood.
- Where you rambled. Depth is going three levels down on one thing, not one level down on ten.
Most people find the same two or three habits every time, which makes them fixable.
Mock interviews, run properly
A mock interview where a friend nods along is worth very little. Give the partner a role and a script.
The partner's job: hold the time budget and announce each phase; give vague answers to clarifying questions, as a real interviewer does; interrupt once with a constraint change around minute 20 — "actually, assume ten times the traffic" or "assume the database can only be eventually consistent"; pick the deep-dive component themselves rather than letting the candidate choose; and ask "why not X?" at least twice, where X is a plausible alternative.
That mid-interview constraint change is the highest-value thing a partner can do. It tests adaptation, which no amount of solo practice exercises.
The rubric to hand your partner
Score each axis 1 to 4, where 3 is a solid senior answer and 4 is a staff-level one. The axes are the four from How you are scored, and how level is decided.
| Axis | 1 — needs work | 2 — mid | 3 — senior | 4 — staff |
|---|---|---|---|---|
| Requirements and scoping | Started designing immediately | Asked some questions, did not use the answers | Bounded the problem, computed numbers, used them later | Surfaced a requirement the interviewer had not considered, and scoped it out with a reason |
| High-level design | Boxes with unlabelled arrows | Reasonable structure, thin data model and API | Clear diagram, API and data model, built up rather than dropped in | Design visibly follows from the numbers; alternatives named at each junction |
| Depth on demand | Could not go deeper than the box | One level down on the chosen component | Three levels down, with mechanisms and figures | Went deep and connected it back to the requirements and the failure modes |
| Trade-offs and communication | Presented one design as the answer | Named alternatives when prompted | Named alternatives unprompted, chose with a reason, critiqued own design | Took a constraint change gracefully and reworked the design without restarting |
Have the partner write one sentence of evidence per score. Sixteen out of sixteen is not the goal; seeing the same axis score 2 three sessions running is the finding.
The specialisation courses
Practice keeps what you have. The next step adds to it. Three courses in this track build directly on this one. Each assumes the vocabulary and the framework you now have, so none of them re-teach quorums or estimation.
Machine Learning System Design Interview. For roles where the product is a model: search ranking, recommendations, feed ranking, fraud detection, ad prediction. It adds the parts this course deliberately left out — training and serving as separate systems, feature stores and the training-serving skew problem, offline and online evaluation, and the fact that a model degrades on its own over time while a database does not. Take it if the job description mentions ranking, recommendations, personalisation, or anything measured by a model metric.
Mobile System Design Interview. For client-heavy roles. It inverts the perspective: the constraints are battery, an intermittent network, a device you cannot deploy to on demand, and offline-first behaviour. Sync, conflict resolution, and local storage move from being one lesson of a section to being the whole course. Take it if you are interviewing for iOS or Android roles, or for anything where the client is the hard part.
Generative AI System Design Interview. For roles building on large language and diffusion models: retrieval-augmented generation, agent systems, embedding search, prompt and context management, and the serving economics that make inference cost a first-class design constraint. Take it if the role is explicitly about building with these models.
If you are not sure, note that Choosing your path, in this course's introduction, gives a path table by role, and it applies to the specialisations as well.
Books worth actually reading
Short list, all substantial, none of which will date quickly.
- Designing Data-Intensive Applications, Martin Kleppmann. The single best book in this area. It covers replication, partitioning, transactions, consensus, and stream processing at more depth than this course can, with the same insistence on trade-offs. If you read one thing, read this.
- Database Internals, Alex Petrov. Storage engines and distributed systems from the inside. The right follow-up if Section 8 (Design a Key-Value Store) was the section you enjoyed most.
- Site Reliability Engineering, by engineers at Google and available to read free online. What operating these systems is actually like. It makes the wrap-up phase of an interview much easier, because it gives you real experience to draw on.
- Release It!, Michael Nygard. Failure modes, stability patterns, and the vocabulary of Reliability patterns with war stories attached.
- Streaming Systems, Tyler Akidau and colleagues. Windows, watermarks, and correctness in stream processing, done properly. The natural extension of Section 23 (Ad Click Event Aggregation).
Papers, and why they are worth the effort
A handful of papers underlie a surprising amount of this course, and reading the original is usually clearer than reading summaries of it. Worth seeking out: Amazon's Dynamo paper, which is behind most of Section 8; the Raft consensus paper, written deliberately for understandability and behind the ordering guarantees in Section 29; Google's Bigtable paper for wide-column storage; and Facebook's Gorilla paper for the time-series compression in Section 22.
Engineering blogs
Company engineering blogs are the closest thing to seeing these decisions made in public. The useful ones publish post-mortems and migration write-ups rather than product announcements — a detailed incident report teaches more about failure modes than any textbook chapter.
A practical habit: when a case study in this course interested you, search for how a company that runs one at scale describes it. Reading their account after having designed it yourself is the same protocol as the solo practice earlier in this lesson — produce first, then compare — applied to real systems.
The habit that matters most
The most reliable way to keep improving is not more reading. It is to keep designing systems you do not have to build, weekly, on paper, in 45 minutes, out loud. That habit costs an hour a week and it is the difference between having read this course and being able to perform it.
You have finished the framework, the building blocks, and twenty-five worked systems. What remains is repetition, and repetition is entirely under your control.