Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
News feed: serving, re-ranking and long-term effects
This lesson covers the request path and the re-ranking pass that enforces everything the model cannot. The model scores each post independently; a feed is a sequence, and several requirements only exist at the level of the sequence.
It then turns to the failures that make this problem discussed outside engineering. They unfold over months and are invisible to the system's own metrics, which is why they need measurements, release gates and user controls of their own.
The path
- Retrieval — the candidate set, from two sources unioned: posts from followed accounts (the fan-out design from System Design Interview) and posts from a recommended pool retrieved by a two-tower model over interest embeddings (Section 6's mechanism). ~1,500 candidates.
- Filter — remove already-seen, muted accounts, blocked users, deleted posts, and policy-removed content. Cheap predicates, and they run before anything expensive.
- Feature hydration — one batched call for candidate and cross features.
- Scoring — the multi-task model over all candidates, producing eight probabilities each.
- Combination — the weighted sum from Predicting several things at once.
- Re-ranking — the pass described below.
- Response — 30 posts.
The latency budget
Invented, against a 300 ms p99 target.
| Stage | Budget |
|---|---|
| Request handling, session state | 10 ms |
| Retrieval: followed fan-out + recommended pool, in parallel | 40 ms |
| Filters | 10 ms |
| Feature hydration (~1,500 candidates) | 55 ms |
| Multi-task scoring (~1,500 × 8 heads) | 85 ms |
| Score combination | 5 ms |
| Re-ranking pass | 20 ms |
| Response assembly, media URLs | 20 ms |
| Network and margin | 55 ms |
| Total | 300 ms |
This budget is tight, and the two large items are hydration and scoring. Two levers worth naming: score only the top ~600 candidates by a cheap pre-score when load is high, and truncate the recommended pool before hydration rather than after.
The re-ranking pass
The model scores each post independently. A feed is a sequence, and several requirements are properties of the sequence rather than of any post in it. Those live in re-ranking.
| Rule | Form | Why the model cannot do it |
|---|---|---|
| Author diversity | At most 2 of 30 from one author | A property of the set |
| Content-type mix | At most 12 of 30 videos | A property of the set |
| Topic diversity | At most 6 of 30 from one topic cluster | A property of the set |
| Freshness floor | At least 8 of 30 posted in the last 6 h | Recency competes with predicted engagement |
| Integrity demotion | Multiply score by Section 5's category scores | Policy, and it changes without retraining |
| Ad load | Insert ads at fixed intervals | Separate system, separate auction (Section 8) |
| Already-seen suppression | Remove or heavily demote | Session state, not model input |
| Exploration slots | 1–2 positions for under-shown or random eligible posts | Deliberately anti-optimal in the short term |
| User controls | Muted words, "see less of this" | Must be exact, not learned |
Implement greedily: take the highest-scored post that violates no constraint, add it, update the constraint state, repeat. It is fast and inspectable.
The argument for rules over training objectives is the one from the video recommendation ranking lesson, and it is worth repeating precisely because it sounds like a compromise and is not. These constraints change on product and policy timelines — sometimes weekly, sometimes in response to an incident. A rule is a configuration change deployable in an hour and auditable afterwards. A learned behaviour is a retraining cycle with an uncertain outcome. For anything that must be exactly true — a muted word never appears, a legally removed post never shows — a rule is the only acceptable implementation.
Pagination and stability
Infinite scroll makes this harder than it looks. New posts arrive while a user scrolls, so offset-based pagination shows duplicates or skips items.
The fix, which is a System Design Interview pattern applied here: pin the candidate set and its ordering at the start of a session in a short-lived session cache, and page through that snapshot. Refresh the snapshot on an explicit pull-to-refresh or after a timeout. The user gets a stable feed; new content appears when they ask for it.
Degradation
Feed requests must never fail:
- Model unavailable → rank by a cheap linear scorer over affinity and recency.
- Feature store slow → hydrate a reduced feature set with defaults, using a model trained to tolerate missing features.
- Recommended pool unavailable → serve followed content only. Less interesting, entirely functional.
- Everything down → reverse chronological from the fan-out cache. The baseline from the framing lesson is also the last line of defence, which is a good reason to keep it working.
Filter bubbles and echo chambers
With the feed serving, the remaining failures are the slow ones. They unfold over months, are invisible to the system's own metrics, and are the reason this problem is discussed outside engineering.
The mechanism, precisely: the model predicts engagement; engagement is higher on content resembling what the user already engages with; so the feed narrows; so the training data narrows; so the model becomes more confident about a smaller region. Nothing malfunctions at any step.
The distinction worth drawing: a filter bubble is narrowing driven by the algorithm; an echo chamber is narrowing driven by the user's own choices about whom to follow. Ranking can worsen or partly counteract the second, and a feed showing only followed accounts still produces echo chambers without any ranking at all. Being precise about which one a design change targets is a mark of a careful answer.
Measurement, since standard metrics are blind to it:
| Metric | Detects |
|---|---|
| Topic entropy per user, tracked over months | Narrowing exposure |
| Source diversity — distinct authors and domains seen per week | Concentration onto a few sources |
| Viewpoint diversity on classifiable topics | The politically charged version, and the hardest to measure well |
| Catalogue coverage — share of posts receiving any impressions | Whether most content reaches anyone |
| Novelty rate — impressions from topics the user has not seen recently | Whether anything new gets through |
Set floors on these and treat a breach as a release blocker, exactly as with engagement regressions. A metric with no consequence attached does not change behaviour.
Misinformation and integrity demotion
Feeds distribute claims at speed, and false claims frequently out-perform true ones on engagement because they are written to. That is a structural property, not an occasional failure.
The response, layered:
- Demote rather than remove, except where policy requires removal. Demotion is reversible, proportionate, and errs toward leaving speech up — the same preference for the reversible action under uncertainty as in Section 5.
- Reduce virality specifically. Cap re-share depth for flagged content, or add friction ("you're about to share an article you haven't opened"). Slowing spread is often more effective than reducing rank, because spread is exponential.
- Attach context rather than only reducing distribution — links to authoritative sources on the claim.
- Demote repeat sources, not only individual posts. A domain that repeatedly publishes debunked claims is evidence about its next article.
- Measure prevalence, not takedowns. As in the harmful content metrics lesson: the fraction of views that land on false content, not the count of items actioned.
Name the tension honestly, because an interviewer will: aggressive demotion suppresses legitimate content, particularly on emerging topics where the truth is not yet established. The resolution is not a threshold; it is graduated responses matched to evidence strength, plus an appeals path.
Giving users real control
Controls are cheap to build and frequently built as decoration. The requirement is that they actually change the ranking:
- "See less of this" must apply an immediate, persistent, strong penalty for that author or topic — not a small feature nudged in the next retraining cycle.
- Mute and block must be exact filters, never soft demotions.
- Feed choice — a chronological option, or a followed-only option. Cheap to implement, and it changes the relationship: a user who can leave the ranked feed is choosing to be in it.
- Explanations — "you're seeing this because you follow X" or "because you engaged with similar posts". Attribute to the candidate source, which is honest and available for free, rather than generating a post-hoc rationalisation of a neural score.
- Topic preferences that feed both re-ranking and training.
A control that does nothing is worse than no control, because it converts a user's attempt to fix their feed into evidence that they cannot.
Evaluating long-term effects
The hardest measurement problem in this section, and the one that makes it a senior topic.
A two-week A/B test cannot measure a three-month effect. Worse, the short-term and long-term effects of an engagement-optimising change frequently have opposite signs: it wins for two weeks and loses over six months, and the two-week result is what ships.
Four approaches, and none is sufficient alone:
- Long-run holdout populations. Keep a group on an older or satisfaction-weighted model for months. Expensive, and the only direct measurement of a months-long effect.
- Surrogate metrics validated against long-term outcomes. Find short-term metrics that historically predict long-term retention, and use those as proxies. Requires historical experiments to validate against, and the relationship can change.
- Longitudinal user surveys. Ask the same users periodically. Slow, and the only direct read on satisfaction rather than behaviour.
- Staged rollouts with extended observation. Roll out over months rather than weeks and watch retention through the whole period, with the ability to reverse.
The honest position, and a strong one to state: this is an unsolved problem. Nobody has a reliable way to measure the six-month effect of a ranking change before shipping it. The mitigations are holdouts, surrogates, and caution about changes that move engagement sharply without moving satisfaction. Saying "we do not have a good answer, here is what we do instead" is more credible than claiming a method that works.