Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Video recommendations: framing, the engagement trap and training data
Prompt: "Design the video recommendation system for the home feed."
This is the canonical recommendation problem and the reference implementation of the two-stage architecture the rest of the course reuses. Step 6: serving introduced the pattern generically; this section builds it properly.
This first lesson frames the problem, deals with the metric conversation — where good candidates separate themselves, because the obvious metrics are the ones that damage the product — and sets up the training data, which is a record of what the previous system chose to show.
The division of labour with Section 10
Two sections in this course cover recommendation ranking, deliberately split:
- This section owns two-stage retrieval and ranking — how millions of items become hundreds, and how those hundreds get ordered.
- Section 10 (Personalized News Feed) owns multi-objective optimisation — how you predict several outcomes at once and combine them into one score, and what the weights mean.
The ranking lesson here uses multi-objective scoring and points at Section 10 for the machinery rather than repeating it.
The clarifying questions
- Which surface? The home feed, which must fill a page from nothing, or the up-next panel, which has a strong contextual anchor in the video currently playing. Different candidate pools, different features, different metrics. Assume home feed.
- What are we optimising? Watch time? Satisfaction? Diversity? Creator ecosystem health? The answer determines everything, and "engagement" is not a sufficient answer — the metrics lesson works through why.
- How large is the catalogue, and how much of it is eligible? 800 million videos exist; perhaps 50 million are plausible recommendations for anyone today.
- How many items per request, and what is the latency budget?
- Cold start: how many users are new, and how many videos are less than an hour old?
The constraints we design against
Invented. The product is Nimbus, the video app used in Sections 1 and 4.
| Constraint | Value |
|---|---|
| Users | 400 million daily active |
| Catalogue | 800 million videos; ~50 million eligible for recommendation |
| Feed requests | ~180,000 per second at peak |
| Items per request | 20, with infinite scroll fetching 20 more |
| Latency budget | 250 ms p99 |
| New videos | ~24 million uploaded per day |
| New users | ~1.5 million per day |
Framing it as machine learning
- Input: a user (with history), a context (time, device, country, session so far), and the eligible catalogue.
- Output: an ordered list of 20 video IDs.
- Objective: maximise expected session value — a weighted blend of watch time, explicit positive signals, and satisfaction, minus explicit negative signals.
The task type is candidate generation, then ranking. Two models, two metrics, two latency budgets, built and evaluated separately.
The arithmetic that forces this. Scoring 50 million candidates with a ranking model at even 0.05 ms each is 2,500 seconds per request. The budget is 250 milliseconds. The gap is four orders of magnitude, and no amount of hardware closes it. So a cheap model must reduce 50 million to about a thousand first.
| Candidate generation | Ranking | |
|---|---|---|
| Input scale | ~50,000,000 | ~1,000 |
| Cost per item | Microseconds | ~0.1 ms |
| Features | Few, precomputable, no user-item crosses | Many, including user-item crosses |
| Metric | recall@1000 | nDCG@20, and online watch time |
| Budget | ~35 ms | ~70 ms |
| Failure mode | Loses the good video permanently | Puts the good video at rank 14 |
The baseline first
Two baselines, and both are worth building.
Popularity. Rank by views in the last 48 hours, with a light penalty for videos the user has already seen. No personalisation whatsoever. It is a surprisingly high bar — popular things are popular because many people like them — and it is the number every personalised model must beat.
Followed creators, reverse chronological. For users who follow anyone, show the newest videos from creators they follow. Zero machine learning, high precision, no cold-start problem for the user, and it fails only in that it cannot fill a feed for users who follow few people and cannot surface anything new.
A blend of these two is a shipping-quality product for a small platform, and it is the correct first system. It also produces the impression and engagement logs that everything afterwards trains on — you cannot train a recommender without a recommender, and this is how you break that circle.
Metrics
The metric conversation in this problem is where good candidates separate themselves, because the obvious metrics are the ones that damage the product.
Offline metrics
Offline evaluation replays logged sessions: given what a user was shown and what they did, would the new model have ranked the engaged item higher?
| Metric | Stage | What it says |
|---|---|---|
| recall@1000 | Candidate generation | Did the item the user actually watched appear in the candidate set? |
| nDCG@20 | Ranking | Is the watched item ranked high, weighted by position |
| precision@k | Ranking | Simpler, coarser; useful for a quick read |
| Predicted-versus-actual watch time correlation | Ranking | Whether the watch-time head is any good, not only whether the ordering is |
All of these are computed on data produced by the current model, which means they are all biased in the way the online metrics lesson described. Offline improvement is a hypothesis, not a result.
Online metrics
| Metric | Layer | Reading |
|---|---|---|
| Click-through rate | Immediate | Are thumbnails and titles working? Weakly related to value |
| Watch time per session | Session | The usual headline metric |
| Completion rate | Session | Did people finish what they started? |
| Session length | Session | Time in the app |
| Next-day return rate | Retention | Did today's feed make them want to come back? |
| 7-day and 28-day active | Retention | The long-run measure, and the slowest |
| Guard: explicit negatives | Immediate | "Not interested", hide, unfollow, report — per thousand impressions |
| Guard: survey satisfaction | Weekly | A sampled in-product question: "was this feed worth your time?" |
The engagement trap, worked through
This is the central metric lesson of this problem, and it belongs in the body of the design.
Optimise watch time alone and the system will find the content that holds attention regardless of whether people are glad it did. In practice that means: longer videos over shorter ones regardless of quality; emotionally activating content over calm content; autoplay chains that are effortless to continue; and content that exploits whatever the user has previously failed to stop watching.
An invented but representative sequence. Nimbus ships a ranker optimising predicted watch time. Over eight weeks:
- Watch time per session rises from 24 to 29 minutes. The launch is declared a success.
- "Not interested" taps per thousand impressions rise from 3.1 to 5.8.
- Weekly survey satisfaction — "was your time on Nimbus today well spent?" — falls from 71% to 64%.
- 28-day active users fall by 0.9%, which takes three months to become statistically visible and is at first attributed to seasonality.
The engagement metric moved in the right direction. The product got worse. By the time retention showed it, the model had been retrained six times on data it had generated, and the catalogue's engagement distribution had shifted to match.
What to do instead
Four practical mechanisms, all of which fit in the design:
- Predict satisfaction, not only engagement. Explicit signals — likes, "not interested", surveys, subscriptions after watching — are sparse but unbiased in a way watch time is not. Predict them as separate heads and include them in the score (see Predicting several things at once).
- Penalise explicit negatives heavily. A "not interested" tap is a strong, deliberate, costly-to-give signal. Weight it far above its frequency would suggest.
- Cap or discount watch time. The marginal value of the 40th minute in a session is not the same as the 4th. A concave transform of watch time reduces the incentive toward endless chains.
- Hold out a population. Keep a small user group on a satisfaction-weighted model and compare long-run retention against the engagement-weighted majority. This is the only way to measure effects that unfold over months.
The offline–online gap here
Expect it, and name the reason. Offline evaluation ranks items the user was already shown. A new model whose value is surfacing items the old model never showed scores worse offline, because the items it would have surfaced have no logged outcome. The models that look best offline are frequently the ones that most resemble the current model.
The mitigations are the ones from the online metrics lesson: propensity logging, a randomised holdout, and treating the offline metric as a filter for what to test rather than as a verdict.
Data and features
Everything the model learns comes from implicit feedback, and implicit feedback is a record of what the previous system chose to show.
Implicit feedback and its three biases
Position bias. Item at rank 1 gets far more engagement than the same item at rank 15. Defined and treated in the video search data and features lesson; the corrections are the same here — position as a training feature set to a constant at serving, inverse propensity weighting, or a randomised slice.
Popularity bias. Popular videos get recommended, which makes them more popular. Two mechanisms compound: the model has more data about them so its predictions are more confident, and confident predictions win ranking. The consequence is a catalogue where a small fraction of videos take the great majority of impressions, and the long tail is invisible regardless of quality. Mitigate by down-weighting training examples by item popularity, and by an explicit exploration budget (see Monitoring and follow-ups).
A non-click is not a negative. The user did not skip item 12; they never scrolled to it. Treating unviewed items as negatives teaches the model that everything below the fold is bad.
Getting negatives right
This is the highest-leverage data decision in this case study, and the two stages need different negatives — a point that is easy to miss.
For ranking, negatives come from the same slate: items that were displayed on screen and not engaged with. Viewport logging is required to know what was actually visible. This is the honest comparison, because ranking's job is to order a slate.
For candidate generation, in-slate negatives are wrong. The retrieval model must distinguish good candidates from the entire catalogue, and it will never see the catalogue if its negatives all come from what the old model already selected. Use random catalogue items, sampled with probability proportional to popularity (so the model is not wasted on items nobody would ever see), plus mined hard negatives — items the current retrieval model ranks highly and that were never engaged with.
Getting this distinction right is worth stating explicitly in an interview. It is a common and costly mistake to train a retrieval model on ranking's negatives, and the symptom is a retrieval stage that never surfaces anything new.
Graded positives
Not all engagement is equal. A hierarchy, and the weight each gets as a training target:
| Signal | Strength | Note |
|---|---|---|
| Subscribed to the creator after watching | Very strong | Rare, and the clearest satisfaction signal available |
| Watched to completion, or > 80% | Strong | Scale by video length — 80% of 12 seconds is weak |
| Liked or shared | Strong | Social act as much as a quality judgement |
| Watched > 30 s | Medium | The usual definition of a real view |
| Clicked, watched < 10 s | Negative | The clickbait signature: attracted, then disappointed |
| Marked "not interested" | Strong negative | Deliberate and costly to give |
| Impression, no click | Weak negative | Only if it was actually on screen |
The fifth row is the important one. A click followed by an immediate exit is worse than no click at all, because the user's time was wasted and the thumbnail misled them. Treating it as a positive because a click occurred is the mechanism by which a system learns to produce clickbait.
Features
User features — long-term topic and creator affinity vectors, watch history summary, demographics where available and permitted, subscription list, language, and behavioural statistics (average session length, typical watch depth).
Item features — creator, topic and category, language, duration, age, quality signals, and aggregate engagement statistics by cohort. Engagement statistics carry most of the signal and cause most of the leakage risk.
Context features — time of day, day of week, device (a phone on a commute is not a TV on a Sunday evening), connection quality, and the session so far. Session context is powerful and underused: what someone has watched in the last ten minutes predicts the next watch better than anything in their long-term history.
Cross features — the user's affinity for this item's topic, whether they follow this creator, how many videos from this creator they finished, time since they last watched this topic. These carry the most signal per feature and are only available at ranking time, because they must be computed per user-item pair. That is the concrete reason the two stages exist and it is worth naming exactly here.
The leakage trap in this problem
An item's "engagement rate over the last 7 days" computed with a daily batch job includes the outcome of the impression being predicted, exactly as in Step 3 continued: features and leakage. On this problem the inflation is large, because item-level engagement rate is one of the strongest features available.
The fix is point-in-time correctness: every aggregate feature is computed as of the impression timestamp, using a feature store that stores values with their valid-from times. Say this explicitly — it is the single most common correctness bug in production recommenders.