Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Video recommendations: framing, the engagement trap and training data


Prompt: "Design the video recommendation system for the home feed."

This is the canonical recommendation problem and the reference implementation of the two-stage architecture the rest of the course reuses. Step 6: serving introduced the pattern generically; this section builds it properly.

This first lesson frames the problem, deals with the metric conversation — where good candidates separate themselves, because the obvious metrics are the ones that damage the product — and sets up the training data, which is a record of what the previous system chose to show.

Millions down to ten, in two stagesBillions of videosSeveral cheap retrieversAbout 1,000 candidatesHeavy ranking modelRe-rank and show ten
Retrieval optimises recall with almost no features; ranking optimises precision on the few candidates it can afford to think about.

The division of labour with Section 10

Two sections in this course cover recommendation ranking, deliberately split:

  • This section owns two-stage retrieval and ranking — how millions of items become hundreds, and how those hundreds get ordered.
  • Section 10 (Personalized News Feed) owns multi-objective optimisation — how you predict several outcomes at once and combine them into one score, and what the weights mean.

The ranking lesson here uses multi-objective scoring and points at Section 10 for the machinery rather than repeating it.

The clarifying questions

  1. Which surface? The home feed, which must fill a page from nothing, or the up-next panel, which has a strong contextual anchor in the video currently playing. Different candidate pools, different features, different metrics. Assume home feed.
  2. What are we optimising? Watch time? Satisfaction? Diversity? Creator ecosystem health? The answer determines everything, and "engagement" is not a sufficient answer — the metrics lesson works through why.
  3. How large is the catalogue, and how much of it is eligible? 800 million videos exist; perhaps 50 million are plausible recommendations for anyone today.
  4. How many items per request, and what is the latency budget?
  5. Cold start: how many users are new, and how many videos are less than an hour old?

The constraints we design against

Invented. The product is Nimbus, the video app used in Sections 1 and 4.

ConstraintValue
Users400 million daily active
Catalogue800 million videos; ~50 million eligible for recommendation
Feed requests~180,000 per second at peak
Items per request20, with infinite scroll fetching 20 more
Latency budget250 ms p99
New videos~24 million uploaded per day
New users~1.5 million per day

Framing it as machine learning

  • Input: a user (with history), a context (time, device, country, session so far), and the eligible catalogue.
  • Output: an ordered list of 20 video IDs.
  • Objective: maximise expected session value — a weighted blend of watch time, explicit positive signals, and satisfaction, minus explicit negative signals.

The task type is candidate generation, then ranking. Two models, two metrics, two latency budgets, built and evaluated separately.

The arithmetic that forces this. Scoring 50 million candidates with a ranking model at even 0.05 ms each is 2,500 seconds per request. The budget is 250 milliseconds. The gap is four orders of magnitude, and no amount of hardware closes it. So a cheap model must reduce 50 million to about a thousand first.

Candidate generationRanking
Input scale~50,000,000~1,000
Cost per itemMicroseconds~0.1 ms
FeaturesFew, precomputable, no user-item crossesMany, including user-item crosses
Metricrecall@1000nDCG@20, and online watch time
Budget~35 ms~70 ms
Failure modeLoses the good video permanentlyPuts the good video at rank 14

The baseline first

Two baselines, and both are worth building.

Popularity. Rank by views in the last 48 hours, with a light penalty for videos the user has already seen. No personalisation whatsoever. It is a surprisingly high bar — popular things are popular because many people like them — and it is the number every personalised model must beat.

Followed creators, reverse chronological. For users who follow anyone, show the newest videos from creators they follow. Zero machine learning, high precision, no cold-start problem for the user, and it fails only in that it cannot fill a feed for users who follow few people and cannot surface anything new.

A blend of these two is a shipping-quality product for a small platform, and it is the correct first system. It also produces the impression and engagement logs that everything afterwards trains on — you cannot train a recommender without a recommender, and this is how you break that circle.

Metrics

The metric conversation in this problem is where good candidates separate themselves, because the obvious metrics are the ones that damage the product.

The engagement trap, in two columnsOptimise watch time alone• Long videos beat short good ones• Outrage and shock hold attention• Feed narrows within a few weeksOptimise a combination• Weighted heads across several actions• Explicit surveys as a slow ground truth• Guard metrics on diversity and regret
The easiest metric to move is the one most likely to make the product worse, which is why the interesting answer is the weighting.

Offline metrics

Offline evaluation replays logged sessions: given what a user was shown and what they did, would the new model have ranked the engaged item higher?

MetricStageWhat it says
recall@1000Candidate generationDid the item the user actually watched appear in the candidate set?
nDCG@20RankingIs the watched item ranked high, weighted by position
precision@kRankingSimpler, coarser; useful for a quick read
Predicted-versus-actual watch time correlationRankingWhether the watch-time head is any good, not only whether the ordering is

All of these are computed on data produced by the current model, which means they are all biased in the way the online metrics lesson described. Offline improvement is a hypothesis, not a result.

Online metrics

MetricLayerReading
Click-through rateImmediateAre thumbnails and titles working? Weakly related to value
Watch time per sessionSessionThe usual headline metric
Completion rateSessionDid people finish what they started?
Session lengthSessionTime in the app
Next-day return rateRetentionDid today's feed make them want to come back?
7-day and 28-day activeRetentionThe long-run measure, and the slowest
Guard: explicit negativesImmediate"Not interested", hide, unfollow, report — per thousand impressions
Guard: survey satisfactionWeeklyA sampled in-product question: "was this feed worth your time?"

The engagement trap, worked through

This is the central metric lesson of this problem, and it belongs in the body of the design.

Optimise watch time alone and the system will find the content that holds attention regardless of whether people are glad it did. In practice that means: longer videos over shorter ones regardless of quality; emotionally activating content over calm content; autoplay chains that are effortless to continue; and content that exploits whatever the user has previously failed to stop watching.

An invented but representative sequence. Nimbus ships a ranker optimising predicted watch time. Over eight weeks:

  • Watch time per session rises from 24 to 29 minutes. The launch is declared a success.
  • "Not interested" taps per thousand impressions rise from 3.1 to 5.8.
  • Weekly survey satisfaction — "was your time on Nimbus today well spent?" — falls from 71% to 64%.
  • 28-day active users fall by 0.9%, which takes three months to become statistically visible and is at first attributed to seasonality.

The engagement metric moved in the right direction. The product got worse. By the time retention showed it, the model had been retrained six times on data it had generated, and the catalogue's engagement distribution had shifted to match.

What to do instead

Four practical mechanisms, all of which fit in the design:

  1. Predict satisfaction, not only engagement. Explicit signals — likes, "not interested", surveys, subscriptions after watching — are sparse but unbiased in a way watch time is not. Predict them as separate heads and include them in the score (see Predicting several things at once).
  2. Penalise explicit negatives heavily. A "not interested" tap is a strong, deliberate, costly-to-give signal. Weight it far above its frequency would suggest.
  3. Cap or discount watch time. The marginal value of the 40th minute in a session is not the same as the 4th. A concave transform of watch time reduces the incentive toward endless chains.
  4. Hold out a population. Keep a small user group on a satisfaction-weighted model and compare long-run retention against the engagement-weighted majority. This is the only way to measure effects that unfold over months.

The offline–online gap here

Expect it, and name the reason. Offline evaluation ranks items the user was already shown. A new model whose value is surfacing items the old model never showed scores worse offline, because the items it would have surfaced have no logged outcome. The models that look best offline are frequently the ones that most resemble the current model.

The mitigations are the ones from the online metrics lesson: propensity logging, a randomised holdout, and treating the offline metric as a filter for what to test rather than as a verdict.

Data and features

Everything the model learns comes from implicit feedback, and implicit feedback is a record of what the previous system chose to show.

Four biases inside one impression logImpression logPosition in the feedThumbnail presentationPopularity of the itemNever shown, no label
Implicit feedback records what the previous ranker chose to show, so training on it teaches the new model to imitate the old one.

Implicit feedback and its three biases

Position bias. Item at rank 1 gets far more engagement than the same item at rank 15. Defined and treated in the video search data and features lesson; the corrections are the same here — position as a training feature set to a constant at serving, inverse propensity weighting, or a randomised slice.

Popularity bias. Popular videos get recommended, which makes them more popular. Two mechanisms compound: the model has more data about them so its predictions are more confident, and confident predictions win ranking. The consequence is a catalogue where a small fraction of videos take the great majority of impressions, and the long tail is invisible regardless of quality. Mitigate by down-weighting training examples by item popularity, and by an explicit exploration budget (see Monitoring and follow-ups).

A non-click is not a negative. The user did not skip item 12; they never scrolled to it. Treating unviewed items as negatives teaches the model that everything below the fold is bad.

Getting negatives right

This is the highest-leverage data decision in this case study, and the two stages need different negatives — a point that is easy to miss.

For ranking, negatives come from the same slate: items that were displayed on screen and not engaged with. Viewport logging is required to know what was actually visible. This is the honest comparison, because ranking's job is to order a slate.

For candidate generation, in-slate negatives are wrong. The retrieval model must distinguish good candidates from the entire catalogue, and it will never see the catalogue if its negatives all come from what the old model already selected. Use random catalogue items, sampled with probability proportional to popularity (so the model is not wasted on items nobody would ever see), plus mined hard negatives — items the current retrieval model ranks highly and that were never engaged with.

Getting this distinction right is worth stating explicitly in an interview. It is a common and costly mistake to train a retrieval model on ranking's negatives, and the symptom is a retrieval stage that never surfaces anything new.

Graded positives

Not all engagement is equal. A hierarchy, and the weight each gets as a training target:

SignalStrengthNote
Subscribed to the creator after watchingVery strongRare, and the clearest satisfaction signal available
Watched to completion, or > 80%StrongScale by video length — 80% of 12 seconds is weak
Liked or sharedStrongSocial act as much as a quality judgement
Watched > 30 sMediumThe usual definition of a real view
Clicked, watched < 10 sNegativeThe clickbait signature: attracted, then disappointed
Marked "not interested"Strong negativeDeliberate and costly to give
Impression, no clickWeak negativeOnly if it was actually on screen

The fifth row is the important one. A click followed by an immediate exit is worse than no click at all, because the user's time was wasted and the thumbnail misled them. Treating it as a positive because a click occurred is the mechanism by which a system learns to produce clickbait.

Features

User features — long-term topic and creator affinity vectors, watch history summary, demographics where available and permitted, subscription list, language, and behavioural statistics (average session length, typical watch depth).

Item features — creator, topic and category, language, duration, age, quality signals, and aggregate engagement statistics by cohort. Engagement statistics carry most of the signal and cause most of the leakage risk.

Context features — time of day, day of week, device (a phone on a commute is not a TV on a Sunday evening), connection quality, and the session so far. Session context is powerful and underused: what someone has watched in the last ten minutes predicts the next watch better than anything in their long-term history.

Cross features — the user's affinity for this item's topic, whether they follow this creator, how many videos from this creator they finished, time since they last watched this topic. These carry the most signal per feature and are only available at ranking time, because they must be computed per user-item pair. That is the concrete reason the two stages exist and it is worth naming exactly here.

The leakage trap in this problem

An item's "engagement rate over the last 7 days" computed with a daily batch job includes the outcome of the impression being predicted, exactly as in Step 3 continued: features and leakage. On this problem the inflation is large, because item-level engagement rate is one of the strongest features available.

The fix is point-in-time correctness: every aggregate feature is computed as of the impression timestamp, using a feature store that stores values with their valid-from times. Say this explicitly — it is the single most common correctness bug in production recommenders.