Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Metrics, labels and features: steps 2 and 3 of the framework
Once the problem is framed, the next ten to thirteen minutes of the round go to two questions: how you will know a model is good, and what it will learn from. Metrics come first because they decide what "better" means; data and features come next because they decide what the model can possibly learn.
Offline metrics let you compare two models on logged data before either touches a user. They are how you decide what to test, not whether to ship. Online metrics decide whether one ships, and the two disagree often enough that the disagreement is itself an interview topic. On the data side, where the labels come from is the question that separates people who have shipped a model from people who have read about them — and leakage is the bug that makes a model look excellent offline and fail completely in production.
Step 2: offline metrics
The four numbers underneath everything
For a binary classifier, every prediction falls into one of four buckets, given a decision threshold:
| Actually positive | Actually negative | |
|---|---|---|
| Predicted positive | True positive (TP) | False positive (FP) |
| Predicted negative | False negative (FN) | True negative (TN) |
- Precision = TP / (TP + FP) — of the things I flagged, how many were right?
- Recall = TP / (TP + FN) — of the things I should have caught, how many did I catch?
- F1 = the harmonic mean of the two — one number when you have no reason to prefer either.
F1 is the metric to be suspicious of. It assumes precision and recall matter equally, and in this course they almost never do. Section 3 blurs faces in street imagery, where a missed face is a privacy incident and an extra blur is a smudge on a wall — recall dominates. Section 5 moderates content, where a false positive silences a person — the balance is different per harm category. Say which error is worse before you pick a metric.
ROC-AUC and PR-AUC, and why the second is the honest one
Both summarise a classifier across all thresholds, so you can compare models without first fixing one.
ROC-AUC plots true positive rate against false positive rate. Read plainly: the probability that a randomly chosen positive is scored above a randomly chosen negative.
PR-AUC plots precision against recall. Its baseline is the positive rate itself, so it moves when the class balance moves.
Here is why the difference matters. Invented numbers, but the shape is real. A harmful-content model scores 1,000,000 posts of which 100 are harmful — a rate of 0.01%.
The model flags 10,000 posts and catches 90 of the 100.
- Recall = 90 / 100 = 90%
- Precision = 90 / 10,000 = 0.9%
- False positive rate = 9,910 / 999,900 ≈ 0.99%
ROC-AUC is computed from recall and false positive rate. Both look excellent: 90% and 1%. The ROC curve will look close to perfect. But 99 out of every 100 posts this model flags are innocent, and if flagging means removal you have wrongly removed 9,910 posts to catch 90.
PR-AUC sees this immediately, because precision is 0.9%. Under heavy imbalance, the enormous negative class makes the false positive rate insensitive — a few thousand mistakes barely register against a million negatives. Precision divides by the number you flagged, so it does register.
Ranking metrics
When the output is an ordered list, precision and recall need a cut-off, and position needs to count.
| Metric | What it measures | What it rewards | What it ignores |
|---|---|---|---|
| Precision@k | Fraction of the top k that are relevant | Getting a clean top slate | Everything below k; order within k |
| Recall@k | Fraction of all relevant items appearing in the top k | Not missing things — the retrieval-stage metric | Order entirely |
| MRR | 1 / rank of the first relevant item, averaged | Getting one right answer to the top | All relevant items after the first |
| mAP | Mean average precision across queries | Getting many relevant items high | Graded relevance — everything is relevant or not |
| nDCG | Discounted gain by position, normalised | Graded relevance and position together | Nothing much — it is the most complete, and the least intuitive |
The practical rule used throughout this course: recall@k for the retrieval stage, nDCG for the ranking stage. Retrieval's job is to not lose the good item; ranking's job is to put it first.
nDCG deserves one plain sentence. Give every result a relevance grade (0 = irrelevant, 3 = perfect). Add the grades up, but divide each by a factor that grows with its position, so a grade-3 item at rank 1 contributes more than the same item at rank 8. Then divide by the score of the best possible ordering, so the result sits between 0 and 1 and is comparable across queries with different numbers of good answers.
Step 2 continued: online metrics and experimentation
Offline metrics rank your candidates. Online metrics decide whether one ships.
The online metric ladder
Online metrics come in layers, and the further down you go, the more they matter and the harder they are to move detectably.
| Layer | Examples | Moves in | Noise |
|---|---|---|---|
| Immediate | Click-through rate, first-click time | Hours | Low |
| Session | Watch time, session length, searches per session | Days | Medium |
| Retention | Next-day return, 7-day active, 28-day active | Weeks | High |
| Business | Revenue per thousand impressions, conversion, subscriptions | Weeks | High |
| Guard | Report rate, hide rate, unsubscribes, support tickets | Days | Medium |
A good interview answer names one from the top of the ladder as the fast signal, one from the middle as the real target, and one guard metric. Three numbers, not one.
A/B testing, in the amount this round requires
Randomise users, not requests. If you randomise requests, the same user sees both variants and the two arms contaminate each other — and for a recommender, the treatment changes what that user's history contains, which changes the control arm's inputs too.
Sample size scales roughly with the metric's variance divided by the square of the effect you want to detect. The practical consequence: small effects on noisy metrics need enormous traffic. Detecting a 1% relative change in a 5%-baseline click-through rate needs on the order of hundreds of thousands of users per arm; detecting a 1% change in 28-day retention needs far more and a month of waiting. These are order-of-magnitude statements, not computed figures — run a real power calculation before committing to an experiment.
Duration must cover at least one full weekly cycle. Behaviour on Tuesday is not behaviour on Saturday.
Novelty and primacy effects distort the first days. A new layout gets clicks because it is new; a changed one loses clicks because users had learned the old one. Both fade. Look at the trend across the experiment, not the first-day delta.
When offline wins and online loses
This happens constantly, and naming the causes is a strong signal.
- Distribution shift from the logging policy. Your offline data was collected by the current model. It only contains items the current model chose to show. A new model that prefers different items is evaluated on a slate it would never have produced.
- Position bias. Logged clicks reflect where an item appeared, not only how good it was. A model trained to predict logged clicks partly learns to predict position.
- Metric mismatch. The offline metric rewards something the online metric does not. Improving nDCG against human relevance grades does not guarantee more watch time.
- The candidate set changed. A better ranker over a worse retrieval stage can score worse than the reverse, and offline evaluation of the ranker alone hides that.
- Latency. A heavier model that is 80 ms slower can lose more engagement to the delay than it gains in relevance. Measure the whole system, not the model.
The mitigations worth naming: log the probability with which each item was shown so you can reweight offline evaluation, keep a small randomised-traffic holdout that produces unbiased data, and run an interleaving experiment — mixing two rankers' results in one list — which detects ranking differences with far less traffic than a split test.
Step 3: data and labels
With the yardstick fixed, turn to what the model learns from. Answer the label question early and in detail.
Three sources of labels
Natural (implicit) labels. The product records the outcome as a by-product of use. A click, a purchase, a watch to 80%, a connection request accepted. Free, plentiful, and immediately biased — you only observe outcomes for items the system chose to show.
Human annotation. People label examples against a rubric. Accurate and expensive. As a planning figure, treat a careful human judgement as costing on the order of tens of seconds to a few minutes each; a million labels is therefore a budget line, not a task. Quality depends on the rubric more than on the annotators, and two annotators will disagree on genuinely ambiguous cases (see data and labels for harmful content detection).
Weak supervision. Cheap, noisy labels generated by rules, heuristics, or existing systems. "Posts removed by the old keyword filter are positives." "Images from the same product listing are similar." Wrong maybe 10–30% of the time, but available in millions. Often the right way to bootstrap, then refine with a small human-labelled set.
| Source | Volume | Cost | Bias risk |
|---|---|---|---|
| Natural | Very high | Near zero | High — reflects what was shown |
| Human | Low | High | Medium — rubric and annotator pool |
| Weak supervision | High | Low | High — inherits the rule's blind spots |
| Self-supervision | Very high | Low | Depends entirely on the augmentation choice |
Self-supervision is the fourth, and Section 2 (Visual Search System) uses it: create pairs from the data itself — two crops of the same image are "similar" by construction — with no labels at all.
Sampling positives and negatives
For a click model, positives are easy: the clicks. Negatives are the hard part, and the choice defines the model.
- Everything not clicked as a negative. Wrong. A user did not click item 14 because they never scrolled to it, not because it was bad. Absence of a click is not a negative.
- Only impressed-and-not-clicked as a negative. Better. It restricts to items the user actually saw.
- Random items from the catalogue as negatives. Necessary for retrieval models, because the retrieval stage must distinguish good items from the whole catalogue, not from the handful the old system already selected.
Hard negatives
This is the idea that most changes model quality, and it recurs in Sections 2, 4, 6 and 9.
A hard negative is an item that is not correct but looks correct. For a visual similarity model on a furniture marketplace: the positive is another photo of the same oak dining chair; an easy negative is a photo of a bicycle; a hard negative is a different oak dining chair with slightly different legs.
Train only on easy negatives and the model learns "chair versus bicycle" — a distinction the product does not need. It will happily rank any chair against any other. Mix in hard negatives and it is forced to encode leg shape, wood grain, and proportion, which is what similarity actually means here.
The practical recipe: a mix, usually a large majority of random or in-batch negatives with a minority of mined hard negatives — items the current model ranks highly but that were not engaged with. Too many hard negatives and training destabilises, because some of them are actually positives you never observed.
Class imbalance
When positives are 0.1% of the data, a model that predicts "negative" always achieves 99.9% accuracy and is worthless. Four strategies, in the order to mention them:
- Downsample the negatives. Keep all positives, keep 1 in 50 negatives. Training gets 50× cheaper. The output probabilities are now wrong and must be corrected — The ad click prediction lesson on data at scale covers the correction, which is a favourite follow-up question.
- Class weights in the loss. Multiply the positive class's loss by a constant. No data is thrown away.
- Focal loss. Reshapes the loss so easy, confidently-correct examples contribute almost nothing, and the gradient concentrates on hard cases. Explained properly in the Street View training lesson and used again in Section 5.
- Change the metric, not the data. Often the imbalance is fine and only the evaluation was misleading. Switch to PR-AUC and per-class recall at a fixed precision.
Step 3 continued: features and leakage
Features turn raw records into model input. Leakage, covered after them, is the most important correctness idea in this course.
Feature types and their treatment
| Type | Examples | Standard treatment |
|---|---|---|
| Numerical | Age of item in hours, price, view count | Scale it — standardise, or log-transform when the distribution has a long tail |
| Low-cardinality categorical | Device type, country, content category | One-hot encoding |
| High-cardinality categorical | User ID, video ID, advertiser ID | A learned embedding, often with hashing (see ad click prediction: features) |
| Text | Title, description, query, transcript | A pretrained text encoder producing a vector; or bag-of-words for a linear baseline |
| Image and video | Thumbnail, frames | A pretrained vision encoder producing a vector (Sections 2, 4, 5) |
| Temporal | Hour of day, day of week | Cyclic encoding — sine and cosine of the angle — so 23:00 sits next to 00:00 |
| Aggregates | Clicks in the last 7 days, average watch time | Windowed counts, and the biggest leakage risk on this list |
Scaling matters for anything gradient-based and matters not at all for tree models, which split on order. That is one reason gradient-boosted trees are so pleasant on tabular data — they need much less feature preparation. Section 7 (Event Recommendation System) leans on this hard.
Embeddings deserve one sentence of explanation: instead of giving a 40-million-value user ID its own column, you give the model a lookup table mapping each ID to, say, a 64-number vector that training adjusts. Similar users drift toward similar vectors. It is compression with learned structure.
Data leakage
Two worked examples, both invented but both common.
The obvious kind. A subscription service builds a churn model. The training table has a column cancellation_reason, non-null only for users who cancelled. The model reaches 0.99 ROC-AUC in an afternoon. In production the field is always null for the users being scored, because they have not cancelled yet. The model is worthless.
The subtle kind. Nimbus builds a "will this video be watched to completion" model. One feature is video_completion_rate_7d, the video's completion rate over the trailing seven days. The training table is built by a nightly job that computes the feature over calendar days. For a view that happened at 09:00 on Tuesday, the seven-day window used in training runs through the end of Tuesday — so it includes the outcome of the very view being predicted, and every other view later that day. Offline ROC-AUC is 0.91. Online it is 0.68.
The fix is point-in-time correctness: every feature value must be computed using only data timestamped before the prediction. This is unglamorous, easy to get wrong, and worth mentioning by name.
Train/serve skew
Leakage is about time. Skew is about two code paths.
Training features are typically computed in a batch job written in one language against a data warehouse. Serving features are computed in a low-latency service written in another language against a different store. The two implementations drift.
An invented but entirely typical case: the training pipeline treats a missing price as 0; the serving path treats it as the category median. Roughly 4% of items have a missing price. The model has learned that price = 0 means "unknown", and at serving time it never sees that signal — it sees a plausible-looking median instead, and confidently scores those items wrong. Nothing errors. No alert fires.
The feature store
A feature store is a service that computes each feature once, from one definition, and serves it to both training and serving.
- Offline store — historical values with timestamps, used to build training sets with point-in-time joins.
- Online store — a low-latency key-value store holding the current value per entity, read during inference.
- One definition — the transformation is written once and materialised into both.
It removes skew by construction and it makes point-in-time correctness a property of the system rather than of the person writing the query. Naming a feature store in the serving section, and saying it exists to prevent skew rather than because it is standard infrastructure, is one of the highest-value sentences available in this round.