Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Metrics, labels and features: steps 2 and 3 of the framework


Once the problem is framed, the next ten to thirteen minutes of the round go to two questions: how you will know a model is good, and what it will learn from. Metrics come first because they decide what "better" means; data and features come next because they decide what the model can possibly learn.

Offline metrics let you compare two models on logged data before either touches a user. They are how you decide what to test, not whether to ship. Online metrics decide whether one ships, and the two disagree often enough that the disagreement is itself an interview topic. On the data side, where the labels come from is the question that separates people who have shipped a model from people who have read about them — and leakage is the bug that makes a model look excellent offline and fail completely in production.

The four numbers under every metricTrue negativeFalsepositiveFalsenegativeTrue positivePredicted 0Predicted 1Actually 0Actually 1Precision reads down the right column; recall reads across the bottom row.
ROC-AUC counts the huge true-negative cell, which is why PR-AUC is the honest metric when positives are rare.

Step 2: offline metrics

The four numbers underneath everything

For a binary classifier, every prediction falls into one of four buckets, given a decision threshold:

Actually positiveActually negative
Predicted positiveTrue positive (TP)False positive (FP)
Predicted negativeFalse negative (FN)True negative (TN)
  • Precision = TP / (TP + FP) — of the things I flagged, how many were right?
  • Recall = TP / (TP + FN) — of the things I should have caught, how many did I catch?
  • F1 = the harmonic mean of the two — one number when you have no reason to prefer either.

F1 is the metric to be suspicious of. It assumes precision and recall matter equally, and in this course they almost never do. Section 3 blurs faces in street imagery, where a missed face is a privacy incident and an extra blur is a smudge on a wall — recall dominates. Section 5 moderates content, where a false positive silences a person — the balance is different per harm category. Say which error is worse before you pick a metric.

ROC-AUC and PR-AUC, and why the second is the honest one

Both summarise a classifier across all thresholds, so you can compare models without first fixing one.

ROC-AUC plots true positive rate against false positive rate. Read plainly: the probability that a randomly chosen positive is scored above a randomly chosen negative.

PR-AUC plots precision against recall. Its baseline is the positive rate itself, so it moves when the class balance moves.

Here is why the difference matters. Invented numbers, but the shape is real. A harmful-content model scores 1,000,000 posts of which 100 are harmful — a rate of 0.01%.

The model flags 10,000 posts and catches 90 of the 100.

  • Recall = 90 / 100 = 90%
  • Precision = 90 / 10,000 = 0.9%
  • False positive rate = 9,910 / 999,900 ≈ 0.99%

ROC-AUC is computed from recall and false positive rate. Both look excellent: 90% and 1%. The ROC curve will look close to perfect. But 99 out of every 100 posts this model flags are innocent, and if flagging means removal you have wrongly removed 9,910 posts to catch 90.

PR-AUC sees this immediately, because precision is 0.9%. Under heavy imbalance, the enormous negative class makes the false positive rate insensitive — a few thousand mistakes barely register against a million negatives. Precision divides by the number you flagged, so it does register.

Ranking metrics

When the output is an ordered list, precision and recall need a cut-off, and position needs to count.

MetricWhat it measuresWhat it rewardsWhat it ignores
Precision@kFraction of the top k that are relevantGetting a clean top slateEverything below k; order within k
Recall@kFraction of all relevant items appearing in the top kNot missing things — the retrieval-stage metricOrder entirely
MRR1 / rank of the first relevant item, averagedGetting one right answer to the topAll relevant items after the first
mAPMean average precision across queriesGetting many relevant items highGraded relevance — everything is relevant or not
nDCGDiscounted gain by position, normalisedGraded relevance and position togetherNothing much — it is the most complete, and the least intuitive

The practical rule used throughout this course: recall@k for the retrieval stage, nDCG for the ranking stage. Retrieval's job is to not lose the good item; ranking's job is to put it first.

nDCG deserves one plain sentence. Give every result a relevance grade (0 = irrelevant, 3 = perfect). Add the grades up, but divide each by a factor that grows with its position, so a grade-3 item at rank 1 contributes more than the same item at rank 8. Then divide by the score of the best possible ordering, so the result sits between 0 and 1 and is comparable across queries with different numbers of good answers.

Step 2 continued: online metrics and experimentation

Offline metrics rank your candidates. Online metrics decide whether one ships.

The online metric ladderClick or immediate actionDwell or completionSession and return rateRetention and revenue
Each rung is closer to the thing you want and slower and noisier to measure, so experiments read the top and guard the bottom.

The online metric ladder

Online metrics come in layers, and the further down you go, the more they matter and the harder they are to move detectably.

LayerExamplesMoves inNoise
ImmediateClick-through rate, first-click timeHoursLow
SessionWatch time, session length, searches per sessionDaysMedium
RetentionNext-day return, 7-day active, 28-day activeWeeksHigh
BusinessRevenue per thousand impressions, conversion, subscriptionsWeeksHigh
GuardReport rate, hide rate, unsubscribes, support ticketsDaysMedium

A good interview answer names one from the top of the ladder as the fast signal, one from the middle as the real target, and one guard metric. Three numbers, not one.

A/B testing, in the amount this round requires

Randomise users, not requests. If you randomise requests, the same user sees both variants and the two arms contaminate each other — and for a recommender, the treatment changes what that user's history contains, which changes the control arm's inputs too.

Sample size scales roughly with the metric's variance divided by the square of the effect you want to detect. The practical consequence: small effects on noisy metrics need enormous traffic. Detecting a 1% relative change in a 5%-baseline click-through rate needs on the order of hundreds of thousands of users per arm; detecting a 1% change in 28-day retention needs far more and a month of waiting. These are order-of-magnitude statements, not computed figures — run a real power calculation before committing to an experiment.

Duration must cover at least one full weekly cycle. Behaviour on Tuesday is not behaviour on Saturday.

Novelty and primacy effects distort the first days. A new layout gets clicks because it is new; a changed one loses clicks because users had learned the old one. Both fade. Look at the trend across the experiment, not the first-day delta.

When offline wins and online loses

This happens constantly, and naming the causes is a strong signal.

  1. Distribution shift from the logging policy. Your offline data was collected by the current model. It only contains items the current model chose to show. A new model that prefers different items is evaluated on a slate it would never have produced.
  2. Position bias. Logged clicks reflect where an item appeared, not only how good it was. A model trained to predict logged clicks partly learns to predict position.
  3. Metric mismatch. The offline metric rewards something the online metric does not. Improving nDCG against human relevance grades does not guarantee more watch time.
  4. The candidate set changed. A better ranker over a worse retrieval stage can score worse than the reverse, and offline evaluation of the ranker alone hides that.
  5. Latency. A heavier model that is 80 ms slower can lose more engagement to the delay than it gains in relevance. Measure the whole system, not the model.

The mitigations worth naming: log the probability with which each item was shown so you can reweight offline evaluation, keep a small randomised-traffic holdout that produces unbiased data, and run an interleaving experiment — mixing two rankers' results in one list — which detects ranking differences with far less traffic than a split test.

Step 3: data and labels

With the yardstick fixed, turn to what the model learns from. Answer the label question early and in detail.

Where a training row comes fromOnelabelled exampleHuman annotationNatural user actionHeuristic labelSampled easy negativeMined hard negative
Naming the label source, and its selection bias, is what separates people who have shipped a model from people who have read about one.

Three sources of labels

Natural (implicit) labels. The product records the outcome as a by-product of use. A click, a purchase, a watch to 80%, a connection request accepted. Free, plentiful, and immediately biased — you only observe outcomes for items the system chose to show.

Human annotation. People label examples against a rubric. Accurate and expensive. As a planning figure, treat a careful human judgement as costing on the order of tens of seconds to a few minutes each; a million labels is therefore a budget line, not a task. Quality depends on the rubric more than on the annotators, and two annotators will disagree on genuinely ambiguous cases (see data and labels for harmful content detection).

Weak supervision. Cheap, noisy labels generated by rules, heuristics, or existing systems. "Posts removed by the old keyword filter are positives." "Images from the same product listing are similar." Wrong maybe 10–30% of the time, but available in millions. Often the right way to bootstrap, then refine with a small human-labelled set.

SourceVolumeCostBias risk
NaturalVery highNear zeroHigh — reflects what was shown
HumanLowHighMedium — rubric and annotator pool
Weak supervisionHighLowHigh — inherits the rule's blind spots
Self-supervisionVery highLowDepends entirely on the augmentation choice

Self-supervision is the fourth, and Section 2 (Visual Search System) uses it: create pairs from the data itself — two crops of the same image are "similar" by construction — with no labels at all.

Sampling positives and negatives

For a click model, positives are easy: the clicks. Negatives are the hard part, and the choice defines the model.

  • Everything not clicked as a negative. Wrong. A user did not click item 14 because they never scrolled to it, not because it was bad. Absence of a click is not a negative.
  • Only impressed-and-not-clicked as a negative. Better. It restricts to items the user actually saw.
  • Random items from the catalogue as negatives. Necessary for retrieval models, because the retrieval stage must distinguish good items from the whole catalogue, not from the handful the old system already selected.

Hard negatives

This is the idea that most changes model quality, and it recurs in Sections 2, 4, 6 and 9.

A hard negative is an item that is not correct but looks correct. For a visual similarity model on a furniture marketplace: the positive is another photo of the same oak dining chair; an easy negative is a photo of a bicycle; a hard negative is a different oak dining chair with slightly different legs.

Train only on easy negatives and the model learns "chair versus bicycle" — a distinction the product does not need. It will happily rank any chair against any other. Mix in hard negatives and it is forced to encode leg shape, wood grain, and proportion, which is what similarity actually means here.

The practical recipe: a mix, usually a large majority of random or in-batch negatives with a minority of mined hard negatives — items the current model ranks highly but that were not engaged with. Too many hard negatives and training destabilises, because some of them are actually positives you never observed.

Class imbalance

When positives are 0.1% of the data, a model that predicts "negative" always achieves 99.9% accuracy and is worthless. Four strategies, in the order to mention them:

  1. Downsample the negatives. Keep all positives, keep 1 in 50 negatives. Training gets 50× cheaper. The output probabilities are now wrong and must be corrected — The ad click prediction lesson on data at scale covers the correction, which is a favourite follow-up question.
  2. Class weights in the loss. Multiply the positive class's loss by a constant. No data is thrown away.
  3. Focal loss. Reshapes the loss so easy, confidently-correct examples contribute almost nothing, and the gradient concentrates on hard cases. Explained properly in the Street View training lesson and used again in Section 5.
  4. Change the metric, not the data. Often the imbalance is fine and only the evaluation was misleading. Switch to PR-AUC and per-class recall at a fixed precision.

Step 3 continued: features and leakage

Features turn raw records into model input. Leakage, covered after them, is the most important correctness idea in this course.

Two ways offline scores lieLeakage• A feature computed after the outcome• Aggregates spanning the label window• Offline AUC of 0.99, real lift of zeroTrain and serve skew• Batch pipelinediffers from the online one• Feature is fresh in training, stale live• Silently wrong for one feature only
A feature store exists so the same code computes the feature in both places, which removes the second bug entirely.

Feature types and their treatment

TypeExamplesStandard treatment
NumericalAge of item in hours, price, view countScale it — standardise, or log-transform when the distribution has a long tail
Low-cardinality categoricalDevice type, country, content categoryOne-hot encoding
High-cardinality categoricalUser ID, video ID, advertiser IDA learned embedding, often with hashing (see ad click prediction: features)
TextTitle, description, query, transcriptA pretrained text encoder producing a vector; or bag-of-words for a linear baseline
Image and videoThumbnail, framesA pretrained vision encoder producing a vector (Sections 2, 4, 5)
TemporalHour of day, day of weekCyclic encoding — sine and cosine of the angle — so 23:00 sits next to 00:00
AggregatesClicks in the last 7 days, average watch timeWindowed counts, and the biggest leakage risk on this list

Scaling matters for anything gradient-based and matters not at all for tree models, which split on order. That is one reason gradient-boosted trees are so pleasant on tabular data — they need much less feature preparation. Section 7 (Event Recommendation System) leans on this hard.

Embeddings deserve one sentence of explanation: instead of giving a 40-million-value user ID its own column, you give the model a lookup table mapping each ID to, say, a 64-number vector that training adjusts. Similar users drift toward similar vectors. It is compression with learned structure.

Data leakage

Two worked examples, both invented but both common.

The obvious kind. A subscription service builds a churn model. The training table has a column cancellation_reason, non-null only for users who cancelled. The model reaches 0.99 ROC-AUC in an afternoon. In production the field is always null for the users being scored, because they have not cancelled yet. The model is worthless.

The subtle kind. Nimbus builds a "will this video be watched to completion" model. One feature is video_completion_rate_7d, the video's completion rate over the trailing seven days. The training table is built by a nightly job that computes the feature over calendar days. For a view that happened at 09:00 on Tuesday, the seven-day window used in training runs through the end of Tuesday — so it includes the outcome of the very view being predicted, and every other view later that day. Offline ROC-AUC is 0.91. Online it is 0.68.

The fix is point-in-time correctness: every feature value must be computed using only data timestamped before the prediction. This is unglamorous, easy to get wrong, and worth mentioning by name.

Train/serve skew

Leakage is about time. Skew is about two code paths.

Training features are typically computed in a batch job written in one language against a data warehouse. Serving features are computed in a low-latency service written in another language against a different store. The two implementations drift.

An invented but entirely typical case: the training pipeline treats a missing price as 0; the serving path treats it as the category median. Roughly 4% of items have a missing price. The model has learned that price = 0 means "unknown", and at serving time it never sees that signal — it sees a plausible-looking median instead, and confidently scores those items wrong. Nothing errors. No alert fires.

The feature store

A feature store is a service that computes each feature once, from one definition, and serves it to both training and serving.

  • Offline store — historical values with timestamps, used to build training sets with point-in-time joins.
  • Online store — a low-latency key-value store holding the current value per entity, read during inference.
  • One definition — the transformation is written once and materialised into both.

It removes skew by construction and it makes point-in-time correctness a property of the system rather than of the person writing the query. Naming a feature store in the serving section, and saying it exists to prevent skew rather than because it is standard infrastructure, is one of the highest-value sentences available in this round.