Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
The ML design round, the seven-step framework and problem framing
The machine learning system design round is two interviews sharing one name, and candidates usually prepare for one of them.
This lesson covers what the round actually scores, the seven-step framework that every case study in this course follows, and the first step of that framework in depth: turning a product wish into a problem a model can be trained on, with an objective you can defend.
The two halves
One half is ordinary system design: queues, caches, sharding, latency budgets, failure modes. That material is covered in System Design Interview, and this course treats it as already known.
The other half is machine learning judgement: what the objective should be, where labels come from, which metric decides success, and how you find out the model has quietly stopped working.
Candidates fail on the half they did not expect. Research-leaning candidates give a strong model discussion and never mention data collection, latency, or drift. Infrastructure-leaning candidates design a beautiful serving path around a model they cannot justify.
What the interviewer is actually scoring
Interview rubrics vary between companies, but the dimensions are consistent. A rough scoring sheet looks like this:
| Dimension | The question behind it | Weight |
|---|---|---|
| Problem framing | Did they turn a vague prompt into an explicit input, output, and objective? | High |
| Data judgement | Do they know where labels come from, and what is wrong with them? | High |
| Metric judgement | Can they pick offline and online metrics and explain disagreement? | High |
| Modelling judgement | Baseline first, then a justified choice — not the newest architecture | Medium |
| Systems judgement | Latency budget, cost, scale, and the serving path | High |
| Monitoring | Drift, feedback loops, retraining | Medium |
| Communication | Structure, thinking aloud, handling redirection | High |
Two things stand out. Data and metrics carry as much weight as the model. And monitoring — the step almost everyone runs out of time for — is scored.
The classic failure
Here are the first sixty seconds of two answers to the same prompt, "design a system to recommend videos". Both candidates are invented.
Candidate A: "I'd use a two-tower model with a transformer encoder on the user side and item embeddings on the other side, trained with sampled softmax…"
Candidate B: "Before I pick anything, let me get the problem straight. Is this the homepage feed or the up-next panel? Those have different candidate pools. What are we optimising — watch time, satisfaction, or something else? Roughly how many videos are in the catalogue, and what's my latency budget for a page load?"
Candidate A has answered a question nobody asked. The architecture might even be right, but the interviewer has no evidence that A knows why, and no evidence A would notice when the data was wrong.
Candidate B has spent sixty seconds and has not named a model. That is correct. The model is step four of seven.
The seven steps
The defence against all three of those notes is a fixed structure. Every case study in this course runs the same seven steps in the same order. Learn the order once and you never again have to decide what to say next.
- Problem framing — clarifying questions, then an explicit input, output, and objective.
- Metrics — offline metrics to select a model, online metrics to decide whether to ship.
- Data and features — sources, labels, feature engineering, and the leakage traps.
- Model — a baseline first, then the real choice, justified against alternatives.
- Training — loss, sampling, imbalance, and the evaluation split.
- Serving — the inference path inside its latency budget.
- Monitoring and follow-ups — drift, feedback loops, retraining, and the extensions this problem attracts.
The order matters more than the content. Steps 1–3 are about the problem; steps 4–5 are about the model; steps 6–7 are about the system that has to live with it. Every step constrains the next. If you have not fixed the objective, you cannot choose a metric. If you have not fixed the metric, you cannot say whether a model is better.
The time budget
A 45-minute round leaves roughly 40 minutes of design time once introductions and closing questions are removed. Divide it like this:
| Step | Minutes | What "done" looks like |
|---|---|---|
| 1 Problem framing | 5 | One sentence stating input, output, objective |
| 2 Metrics | 5 | One offline metric, one online metric, one guard metric |
| 3 Data and features | 8 | Label source named, top features listed, one leakage risk called out |
| 4 Model | 7 | A baseline and one step up, with the reason for the step |
| 5 Training | 5 | Loss, imbalance strategy, and a time-based split |
| 6 Serving | 8 | A drawn diagram with a latency budget that adds up |
| 7 Monitoring | 4 | Two drift signals and a retraining trigger |
Using the budget as a live instrument
The budget is not a plan you make once. It is something you check against the clock. If twenty minutes have gone and you are still discussing labels, you have over-invested in step 3 and will not reach serving.
The recovery is a sentence, said out loud: "I could keep going on labelling, but I want to make sure we get to the serving path — can I come back to this if there's time?" Interviewers read that as time management, not as retreat.
Where to deviate
The order is a default, not a law. Two deviations are common and both are fine if you name them.
If the interviewer opens with "assume we already have a ranking model, focus on serving it", compress steps 1–5 into two minutes and spend the round on step 6.
If the problem is fundamentally about data — harmful content detection, for example, where labelling is the hard part — spend twelve minutes on step 3 and say why you are doing it.
Step 1: framing a product problem as an ML problem
The prompt you are given is a product wish. Your first job is to convert it into something a model can be trained to do.
The conversion
An ML problem is fully specified by three things:
- Input — what the model sees at prediction time, and only what is available then.
- Output — the exact shape of the prediction. A number? A label? A ranked list? Boxes?
- Objective — the quantity being maximised or minimised.
Take the prompt "show better videos". Here are three legitimate conversions of it, using an invented video app called Nimbus as the running example throughout this section.
| Framing | Input | Output | Objective |
|---|---|---|---|
| A | (user, video) pair | Probability the user watches past 30 seconds | Maximise expected watch starts |
| B | user + context | Ranked list of 20 videos from the catalogue | Maximise total watch time in the session |
| C | user + context | Ranked list, scored on watch, like, share, and "not interested" | Maximise a weighted blend of engagement and satisfaction |
All three are defensible. They imply different data, different metrics, and different serving paths. Picking one and saying why is step 1. Not picking one and drifting between them is the most common way a good answer becomes incoherent by minute twenty.
Choosing the task type
Four task types cover most of this course.
- Classification — output is a label or a probability over labels. Is this post harmful? Will this ad be clicked? (Sections 5 and 8.)
- Regression — output is a continuous number. How many minutes will this be watched?
- Ranking — inputs are a query or user plus a candidate set, and the output is an ordering. Most recommendation and search problems. (Sections 4, 6, 7 and 10.)
- Retrieval — output is a small subset pulled from a very large pool, usually by nearest neighbour in an embedding space. (Sections 2 and 9, and the first stage of Sections 4 and 6.)
Two more appear in specific lessons: detection, which outputs boxes and labels over an image (Section 3, Street View blurring), and link prediction on a graph (Section 11, People You May Know).
The choice between ranking and classification is subtle and worth stating clearly. If your scoring model outputs a per-item probability and you sort by it, you have built a classifier and used it as a ranker. That is fine and extremely common. It becomes wrong when the probability itself is used for something — pricing an ad, for instance — because then calibration matters and ranking quality alone does not. Section 8 (Ad Click Prediction) is entirely about that distinction.
The questions you are allowed to ask
Every case study in this course opens with clarifying questions. Four categories cover almost all of them:
- Scope — which surface, which users, which content types?
- Objective — what does the business actually want more of?
- Scale — how many users, items, requests per second, and how much history?
- Constraints — latency budget, cost, freshness, privacy, and regulatory limits.
Ask four to six. Then stop and state your framing. Endless clarification reads as stalling.
When it should not be machine learning
Saying this out loud is a strong signal, not a weak one.
Machine learning is the wrong tool when the rule is known and stable ("block posts containing this exact banned phrase"), when there are no labels and no cheap way to get them, when the volume is too small for a model to beat a heuristic, or when an error is catastrophic and unexplainable errors are unacceptable.
A realistic version: for Nimbus's "trending now" shelf, sorting by view velocity over the last six hours with a small popularity penalty for repeat creators is a rule. It needs no training data and no monitoring beyond a dashboard. Proposing a model for it would be worse engineering, and an interviewer will notice.
Choosing the right objective
The framing above named an objective, and that choice deserves its own scrutiny. The thing you can measure is almost never the thing you want. Every objective in this course is a proxy, and every proxy has a failure mode.
The gap between want and measure
Nimbus wants users to find videos they are glad they watched. Nothing in the logs records gladness. So the team picks something recorded: a click, a watch, a like, minutes watched.
Each substitution introduces a gap, and a model trained hard enough will find and exploit it.
| Proxy objective | What the model learns to produce | The gap |
|---|---|---|
| Click-through rate | Shocking thumbnails and misleading titles | A click measures curiosity, not value |
| Watch time | Long, slow, padded videos | Time spent is not time enjoyed |
| Completion rate | Very short videos | Completion is easy when the video is 8 seconds |
| Likes | Content from creators with mobilised fanbases | Likes are a social act, not a quality rating |
| Session length | Autoplay chains that are hard to leave | Difficulty leaving is not satisfaction |
This is Goodhart's law: when a measure becomes a target, it stops being a good measure. In machine learning systems it is not a philosophical worry. It is the predictable result of optimising a differentiable proxy for millions of steps.
A concrete failure
Invented but representative. Nimbus switches its ranking objective from "will the user click" to "will the user click", weighted by nothing else. Over six weeks:
- Click-through rate on the homepage rises from 8.1% to 10.4% — the launch looks like a win.
- Median watch time per opened video falls from 2:40 to 1:05.
- The share of sessions ending within 30 seconds of the first click rises from 12% to 19%.
- Next-day return rate falls by about 1.5 points.
The model did what it was told. Users clicked more and stayed less. The team shipped a regression and had a dashboard saying otherwise for six weeks.
Guard metrics
A guard metric is chosen to move in the opposite direction if the proxy is being gamed.
- Optimising click-through rate? Guard on post-click watch time and on quick back-outs.
- Optimising watch time? Guard on next-day return and on explicit "not interested" taps.
- Optimising engagement on a feed? Guard on user-reported satisfaction surveys and on hide/report rates.
- Optimising ad revenue? Guard on ad-hide rate and organic session length.
Say the guard metric in the interview at the same moment you name the objective. It takes one clause and it is one of the cheapest ways to sound like someone who has shipped a model.
Combining objectives
When one number cannot capture the goal, predict several and combine them:
score = w1·P(watch) + w2·P(like) + w3·P(share) − w4·P(hide) − w5·P(report)The weights are a product decision, not a statistical one. There is no data-derived value for "how much is one report worth in units of likes". Someone chooses it, and defending that choice is a design conversation.
This is multi-objective ranking. Section 10 (Personalized News Feed) owns it in full — how the heads are trained, how the weights are set and re-tuned, and how negative signals get their scale. Section 6 (Video Recommendation System) uses the same idea for video ranking and points at Section 10 for the machinery.