Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

News feed: serving, re-ranking and long-term effects


This lesson covers the request path and the re-ranking pass that enforces everything the model cannot. The model scores each post independently; a feed is a sequence, and several requirements only exist at the level of the sequence.

It then turns to the failures that make this problem discussed outside engineering. They unfold over months and are invisible to the system's own metrics, which is why they need measurements, release gates and user controls of their own.

What the re-ranking pass adds on topCalibratedmodel scoreIntegrity demotionDiversityand dedupeAuthor variety capStablepage boundarytopbottomEverything above the score is a rule the model is not allowed to learn.
The re-ranker enforces the constraints that must hold on every page, which no per-item score can express on its own.

The path

  1. Retrieval — the candidate set, from two sources unioned: posts from followed accounts (the fan-out design from System Design Interview) and posts from a recommended pool retrieved by a two-tower model over interest embeddings (Section 6's mechanism). ~1,500 candidates.
  2. Filter — remove already-seen, muted accounts, blocked users, deleted posts, and policy-removed content. Cheap predicates, and they run before anything expensive.
  3. Feature hydration — one batched call for candidate and cross features.
  4. Scoring — the multi-task model over all candidates, producing eight probabilities each.
  5. Combination — the weighted sum from Predicting several things at once.
  6. Re-ranking — the pass described below.
  7. Response — 30 posts.

The latency budget

Invented, against a 300 ms p99 target.

StageBudget
Request handling, session state10 ms
Retrieval: followed fan-out + recommended pool, in parallel40 ms
Filters10 ms
Feature hydration (~1,500 candidates)55 ms
Multi-task scoring (~1,500 × 8 heads)85 ms
Score combination5 ms
Re-ranking pass20 ms
Response assembly, media URLs20 ms
Network and margin55 ms
Total300 ms

This budget is tight, and the two large items are hydration and scoring. Two levers worth naming: score only the top ~600 candidates by a cheap pre-score when load is high, and truncate the recommended pool before hydration rather than after.

The re-ranking pass

The model scores each post independently. A feed is a sequence, and several requirements are properties of the sequence rather than of any post in it. Those live in re-ranking.

RuleFormWhy the model cannot do it
Author diversityAt most 2 of 30 from one authorA property of the set
Content-type mixAt most 12 of 30 videosA property of the set
Topic diversityAt most 6 of 30 from one topic clusterA property of the set
Freshness floorAt least 8 of 30 posted in the last 6 hRecency competes with predicted engagement
Integrity demotionMultiply score by Section 5's category scoresPolicy, and it changes without retraining
Ad loadInsert ads at fixed intervalsSeparate system, separate auction (Section 8)
Already-seen suppressionRemove or heavily demoteSession state, not model input
Exploration slots1–2 positions for under-shown or random eligible postsDeliberately anti-optimal in the short term
User controlsMuted words, "see less of this"Must be exact, not learned

Implement greedily: take the highest-scored post that violates no constraint, add it, update the constraint state, repeat. It is fast and inspectable.

The argument for rules over training objectives is the one from the video recommendation ranking lesson, and it is worth repeating precisely because it sounds like a compromise and is not. These constraints change on product and policy timelines — sometimes weekly, sometimes in response to an incident. A rule is a configuration change deployable in an hour and auditable afterwards. A learned behaviour is a retraining cycle with an uncertain outcome. For anything that must be exactly true — a muted word never appears, a legally removed post never shows — a rule is the only acceptable implementation.

Pagination and stability

Infinite scroll makes this harder than it looks. New posts arrive while a user scrolls, so offset-based pagination shows duplicates or skips items.

The fix, which is a System Design Interview pattern applied here: pin the candidate set and its ordering at the start of a session in a short-lived session cache, and page through that snapshot. Refresh the snapshot on an explicit pull-to-refresh or after a timeout. The user gets a stable feed; new content appears when they ask for it.

Degradation

Feed requests must never fail:

  • Model unavailable → rank by a cheap linear scorer over affinity and recency.
  • Feature store slow → hydrate a reduced feature set with defaults, using a model trained to tolerate missing features.
  • Recommended pool unavailable → serve followed content only. Less interesting, entirely functional.
  • Everything down → reverse chronological from the fan-out cache. The baseline from the framing lesson is also the last line of defence, which is a good reason to keep it working.

Filter bubbles and echo chambers

With the feed serving, the remaining failures are the slow ones. They unfold over months, are invisible to the system's own metrics, and are the reason this problem is discussed outside engineering.

Two clocks, and only one has a dashboardWhat daily metrics see• Sessions, interactions, time spent• Moves within hours of a launch• All healthy while the feed narrowsWhat unfolds over months• Narrowing of sources a user sees• Misinformation reach before demotion• Users who quietly stop returning
Long-horizon holdout groups exist because a metric measured this week cannot detect a harm that takes a year.

The mechanism, precisely: the model predicts engagement; engagement is higher on content resembling what the user already engages with; so the feed narrows; so the training data narrows; so the model becomes more confident about a smaller region. Nothing malfunctions at any step.

The distinction worth drawing: a filter bubble is narrowing driven by the algorithm; an echo chamber is narrowing driven by the user's own choices about whom to follow. Ranking can worsen or partly counteract the second, and a feed showing only followed accounts still produces echo chambers without any ranking at all. Being precise about which one a design change targets is a mark of a careful answer.

Measurement, since standard metrics are blind to it:

MetricDetects
Topic entropy per user, tracked over monthsNarrowing exposure
Source diversity — distinct authors and domains seen per weekConcentration onto a few sources
Viewpoint diversity on classifiable topicsThe politically charged version, and the hardest to measure well
Catalogue coverage — share of posts receiving any impressionsWhether most content reaches anyone
Novelty rate — impressions from topics the user has not seen recentlyWhether anything new gets through

Set floors on these and treat a breach as a release blocker, exactly as with engagement regressions. A metric with no consequence attached does not change behaviour.

Misinformation and integrity demotion

Feeds distribute claims at speed, and false claims frequently out-perform true ones on engagement because they are written to. That is a structural property, not an occasional failure.

The response, layered:

  1. Demote rather than remove, except where policy requires removal. Demotion is reversible, proportionate, and errs toward leaving speech up — the same preference for the reversible action under uncertainty as in Section 5.
  2. Reduce virality specifically. Cap re-share depth for flagged content, or add friction ("you're about to share an article you haven't opened"). Slowing spread is often more effective than reducing rank, because spread is exponential.
  3. Attach context rather than only reducing distribution — links to authoritative sources on the claim.
  4. Demote repeat sources, not only individual posts. A domain that repeatedly publishes debunked claims is evidence about its next article.
  5. Measure prevalence, not takedowns. As in the harmful content metrics lesson: the fraction of views that land on false content, not the count of items actioned.

Name the tension honestly, because an interviewer will: aggressive demotion suppresses legitimate content, particularly on emerging topics where the truth is not yet established. The resolution is not a threshold; it is graduated responses matched to evidence strength, plus an appeals path.

Giving users real control

Controls are cheap to build and frequently built as decoration. The requirement is that they actually change the ranking:

  • "See less of this" must apply an immediate, persistent, strong penalty for that author or topic — not a small feature nudged in the next retraining cycle.
  • Mute and block must be exact filters, never soft demotions.
  • Feed choice — a chronological option, or a followed-only option. Cheap to implement, and it changes the relationship: a user who can leave the ranked feed is choosing to be in it.
  • Explanations — "you're seeing this because you follow X" or "because you engaged with similar posts". Attribute to the candidate source, which is honest and available for free, rather than generating a post-hoc rationalisation of a neural score.
  • Topic preferences that feed both re-ranking and training.

A control that does nothing is worse than no control, because it converts a user's attempt to fix their feed into evidence that they cannot.

Evaluating long-term effects

The hardest measurement problem in this section, and the one that makes it a senior topic.

A two-week A/B test cannot measure a three-month effect. Worse, the short-term and long-term effects of an engagement-optimising change frequently have opposite signs: it wins for two weeks and loses over six months, and the two-week result is what ships.

Four approaches, and none is sufficient alone:

  1. Long-run holdout populations. Keep a group on an older or satisfaction-weighted model for months. Expensive, and the only direct measurement of a months-long effect.
  2. Surrogate metrics validated against long-term outcomes. Find short-term metrics that historically predict long-term retention, and use those as proxies. Requires historical experiments to validate against, and the relationship can change.
  3. Longitudinal user surveys. Ask the same users periodically. Slow, and the only direct read on satisfaction rather than behaviour.
  4. Staged rollouts with extended observation. Roll out over months rather than weeks and watch retention through the whole period, with the ability to reverse.

The honest position, and a strong one to state: this is an unsolved problem. Nobody has a reliable way to measure the six-month effect of a ranking change before shipping it. The mitigations are holdouts, surrogates, and caution about changes that move engagement sharply without moving satisfaction. Saying "we do not have a good answer, here is what we do instead" is more credible than claiming a method that works.