Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

The ML design round, the seven-step framework and problem framing


The machine learning system design round is two interviews sharing one name, and candidates usually prepare for one of them.

This lesson covers what the round actually scores, the seven-step framework that every case study in this course follows, and the first step of that framework in depth: turning a product wish into a problem a model can be trained on, with an objective you can defend.

Two interviews sharing one nameThe framework half• Framing a wish as a trainable task• Metrics, data, model, serving, in order• Naming trade-offs before choosingThe depth half• Where the labels actually come from• Latency budgets that constrain the model• Failure modes you have seen happen
Candidates rehearse the framework and get scored on the depth they hang off it.

The two halves

One half is ordinary system design: queues, caches, sharding, latency budgets, failure modes. That material is covered in System Design Interview, and this course treats it as already known.

The other half is machine learning judgement: what the objective should be, where labels come from, which metric decides success, and how you find out the model has quietly stopped working.

Candidates fail on the half they did not expect. Research-leaning candidates give a strong model discussion and never mention data collection, latency, or drift. Infrastructure-leaning candidates design a beautiful serving path around a model they cannot justify.

What the interviewer is actually scoring

Interview rubrics vary between companies, but the dimensions are consistent. A rough scoring sheet looks like this:

DimensionThe question behind itWeight
Problem framingDid they turn a vague prompt into an explicit input, output, and objective?High
Data judgementDo they know where labels come from, and what is wrong with them?High
Metric judgementCan they pick offline and online metrics and explain disagreement?High
Modelling judgementBaseline first, then a justified choice — not the newest architectureMedium
Systems judgementLatency budget, cost, scale, and the serving pathHigh
MonitoringDrift, feedback loops, retrainingMedium
CommunicationStructure, thinking aloud, handling redirectionHigh

Two things stand out. Data and metrics carry as much weight as the model. And monitoring — the step almost everyone runs out of time for — is scored.

The classic failure

Here are the first sixty seconds of two answers to the same prompt, "design a system to recommend videos". Both candidates are invented.

Candidate A: "I'd use a two-tower model with a transformer encoder on the user side and item embeddings on the other side, trained with sampled softmax…"

Candidate B: "Before I pick anything, let me get the problem straight. Is this the homepage feed or the up-next panel? Those have different candidate pools. What are we optimising — watch time, satisfaction, or something else? Roughly how many videos are in the catalogue, and what's my latency budget for a page load?"

Candidate A has answered a question nobody asked. The architecture might even be right, but the interviewer has no evidence that A knows why, and no evidence A would notice when the data was wrong.

Candidate B has spent sixty seconds and has not named a model. That is correct. The model is step four of seven.

The seven steps

The defence against all three of those notes is a fixed structure. Every case study in this course runs the same seven steps in the same order. Learn the order once and you never again have to decide what to say next.

  1. Problem framing — clarifying questions, then an explicit input, output, and objective.
  2. Metrics — offline metrics to select a model, online metrics to decide whether to ship.
  3. Data and features — sources, labels, feature engineering, and the leakage traps.
  4. Model — a baseline first, then the real choice, justified against alternatives.
  5. Training — loss, sampling, imbalance, and the evaluation split.
  6. Serving — the inference path inside its latency budget.
  7. Monitoring and follow-ups — drift, feedback loops, retraining, and the extensions this problem attracts.

The order matters more than the content. Steps 1–3 are about the problem; steps 4–5 are about the model; steps 6–7 are about the system that has to live with it. Every step constrains the next. If you have not fixed the objective, you cannot choose a metric. If you have not fixed the metric, you cannot say whether a model is better.

The time budget

A 45-minute round leaves roughly 40 minutes of design time once introductions and closing questions are removed. Divide it like this:

StepMinutesWhat "done" looks like
1 Problem framing5One sentence stating input, output, objective
2 Metrics5One offline metric, one online metric, one guard metric
3 Data and features8Label source named, top features listed, one leakage risk called out
4 Model7A baseline and one step up, with the reason for the step
5 Training5Loss, imbalance strategy, and a time-based split
6 Serving8A drawn diagram with a latency budget that adds up
7 Monitoring4Two drift signals and a retraining trigger
Clarify the problem5mFrame as ML5mwhat is the label?Data7msources, labels, leakageFeatures6mModel6mEvaluation8moffline and onlineServing + monitoring8m45 minutesframing is where candidates lose the round:name the label and the prediction unit out loudevaluation and serving are half theinterview and the half that gets skippedData and evaluation together outweigh model choice — which is the opposite of how most candidates spend the time.
Model selection gets six minutes; the interviewer is far more interested in the label definition and the eval plan.

Using the budget as a live instrument

The budget is not a plan you make once. It is something you check against the clock. If twenty minutes have gone and you are still discussing labels, you have over-invested in step 3 and will not reach serving.

The recovery is a sentence, said out loud: "I could keep going on labelling, but I want to make sure we get to the serving path — can I come back to this if there's time?" Interviewers read that as time management, not as retreat.

Where to deviate

The order is a default, not a law. Two deviations are common and both are fine if you name them.

If the interviewer opens with "assume we already have a ranking model, focus on serving it", compress steps 1–5 into two minutes and spend the round on step 6.

If the problem is fundamentally about data — harmful content detection, for example, where labelling is the hard part — spend twelve minutes on step 3 and say why you are doing it.

Step 1: framing a product problem as an ML problem

The prompt you are given is a product wish. Your first job is to convert it into something a model can be trained to do.

From product wish to trainable taskProduct wishWho andwhat changesTask typeInput and outputThe labelIf you cannot name the label, it is not yet an ML problem.
The framing is finished only when one row of training data can be written down in full.

The conversion

An ML problem is fully specified by three things:

  • Input — what the model sees at prediction time, and only what is available then.
  • Output — the exact shape of the prediction. A number? A label? A ranked list? Boxes?
  • Objective — the quantity being maximised or minimised.

Take the prompt "show better videos". Here are three legitimate conversions of it, using an invented video app called Nimbus as the running example throughout this section.

FramingInputOutputObjective
A(user, video) pairProbability the user watches past 30 secondsMaximise expected watch starts
Buser + contextRanked list of 20 videos from the catalogueMaximise total watch time in the session
Cuser + contextRanked list, scored on watch, like, share, and "not interested"Maximise a weighted blend of engagement and satisfaction

All three are defensible. They imply different data, different metrics, and different serving paths. Picking one and saying why is step 1. Not picking one and drifting between them is the most common way a good answer becomes incoherent by minute twenty.

Choosing the task type

Four task types cover most of this course.

  • Classification — output is a label or a probability over labels. Is this post harmful? Will this ad be clicked? (Sections 5 and 8.)
  • Regression — output is a continuous number. How many minutes will this be watched?
  • Ranking — inputs are a query or user plus a candidate set, and the output is an ordering. Most recommendation and search problems. (Sections 4, 6, 7 and 10.)
  • Retrieval — output is a small subset pulled from a very large pool, usually by nearest neighbour in an embedding space. (Sections 2 and 9, and the first stage of Sections 4 and 6.)

Two more appear in specific lessons: detection, which outputs boxes and labels over an image (Section 3, Street View blurring), and link prediction on a graph (Section 11, People You May Know).

The choice between ranking and classification is subtle and worth stating clearly. If your scoring model outputs a per-item probability and you sort by it, you have built a classifier and used it as a ranker. That is fine and extremely common. It becomes wrong when the probability itself is used for something — pricing an ad, for instance — because then calibration matters and ranking quality alone does not. Section 8 (Ad Click Prediction) is entirely about that distinction.

The questions you are allowed to ask

Every case study in this course opens with clarifying questions. Four categories cover almost all of them:

  1. Scope — which surface, which users, which content types?
  2. Objective — what does the business actually want more of?
  3. Scale — how many users, items, requests per second, and how much history?
  4. Constraints — latency budget, cost, freshness, privacy, and regulatory limits.

Ask four to six. Then stop and state your framing. Endless clarification reads as stalling.

When it should not be machine learning

Saying this out loud is a strong signal, not a weak one.

Machine learning is the wrong tool when the rule is known and stable ("block posts containing this exact banned phrase"), when there are no labels and no cheap way to get them, when the volume is too small for a model to beat a heuristic, or when an error is catastrophic and unexplainable errors are unacceptable.

A realistic version: for Nimbus's "trending now" shelf, sorting by view velocity over the last six hours with a small popularity penalty for repeat creators is a rule. It needs no training data and no monitoring beyond a dashboard. Proposing a model for it would be worse engineering, and an interviewer will notice.

Choosing the right objective

The framing above named an objective, and that choice deserves its own scrutiny. The thing you can measure is almost never the thing you want. Every objective in this course is a proxy, and every proxy has a failure mode.

The proxy and the thing you wantedWhat you can measure• Clicks, watch time, session length• Available in the log tonight• Moves quickly under optimisationWhat you actually want• Users who come back next month• Content they are glad they saw• Visible only over long horizons
Every objective here is a proxy, so the design question is which guard metric catches its failure.

The gap between want and measure

Nimbus wants users to find videos they are glad they watched. Nothing in the logs records gladness. So the team picks something recorded: a click, a watch, a like, minutes watched.

Each substitution introduces a gap, and a model trained hard enough will find and exploit it.

Proxy objectiveWhat the model learns to produceThe gap
Click-through rateShocking thumbnails and misleading titlesA click measures curiosity, not value
Watch timeLong, slow, padded videosTime spent is not time enjoyed
Completion rateVery short videosCompletion is easy when the video is 8 seconds
LikesContent from creators with mobilised fanbasesLikes are a social act, not a quality rating
Session lengthAutoplay chains that are hard to leaveDifficulty leaving is not satisfaction

This is Goodhart's law: when a measure becomes a target, it stops being a good measure. In machine learning systems it is not a philosophical worry. It is the predictable result of optimising a differentiable proxy for millions of steps.

A concrete failure

Invented but representative. Nimbus switches its ranking objective from "will the user click" to "will the user click", weighted by nothing else. Over six weeks:

  • Click-through rate on the homepage rises from 8.1% to 10.4% — the launch looks like a win.
  • Median watch time per opened video falls from 2:40 to 1:05.
  • The share of sessions ending within 30 seconds of the first click rises from 12% to 19%.
  • Next-day return rate falls by about 1.5 points.

The model did what it was told. Users clicked more and stayed less. The team shipped a regression and had a dashboard saying otherwise for six weeks.

Guard metrics

A guard metric is chosen to move in the opposite direction if the proxy is being gamed.

  • Optimising click-through rate? Guard on post-click watch time and on quick back-outs.
  • Optimising watch time? Guard on next-day return and on explicit "not interested" taps.
  • Optimising engagement on a feed? Guard on user-reported satisfaction surveys and on hide/report rates.
  • Optimising ad revenue? Guard on ad-hide rate and organic session length.

Say the guard metric in the interview at the same moment you name the objective. It takes one clause and it is one of the cheapest ways to sound like someone who has shipped a model.

Combining objectives

When one number cannot capture the goal, predict several and combine them:

Text
score = w1·P(watch) + w2·P(like) + w3·P(share) − w4·P(hide) − w5·P(report)

The weights are a product decision, not a statistical one. There is no data-derived value for "how much is one report worth in units of likes". Someone chooses it, and defending that choice is a design conversation.

This is multi-objective ranking. Section 10 (Personalized News Feed) owns it in full — how the heads are trained, how the weights are set and re-tuned, and how negative signals get their scale. Section 6 (Video Recommendation System) uses the same idea for video ranking and points at Section 10 for the machinery.