Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

News feed: framing, costly signals and a biased impression log


Prompt: "Design the ranking system for a personalised news feed."

This section owns multi-objective optimisation — predicting several outcomes at once and combining them into a single score. It is the technique Section 6 used and deferred, and it is the heart of how every large feed is actually ranked.

This first lesson sets up why a single objective will not do. It frames the problem, makes the case that no single metric captures a good feed, and builds the training data from an impression log that describes a feed the system already chose.

Inventory to page, per requestRecent inventoryCandidatesourcesMulti-headrankerCombineinto one scoreRe-rank the pageInventory is what friends and pages posted since the last visit.
The candidate set is bounded by connections rather than by a catalogue, which makes retrieval cheap and ranking the whole problem.

The relationship to the news feed design in System Design Interview

This section and the news feed case study in System Design Interview design the same product and solve opposite halves of it. Saying so out loud in an interview is a genuinely good move, because it shows you know which problem you are being asked about.

System Design Interview: news feedThis section
The questionHow does a post get from an author to a follower's feed?Given the posts that could be shown, in what order?
The core trade-offFan-out on write versus fan-out on readWhich objectives, and with what weights
The hard partCelebrity accounts with 40 million followersObjectives that pull against each other
The outputA set of candidate post IDs, unorderedAn ordered list of 30
Failure modeFeed loads slowly, or misses postsFeed loads fast and is not worth reading

The system design version produces the candidate set this section ranks. A candidate who has done both can say: "the fan-out design gives me a few hundred to a couple of thousand candidates per request; my job here starts from there." That sentence connects two rounds of preparation and takes four seconds.

The clarifying questions

  1. Which content sources? Only posts from accounts the user follows, or also recommended content from accounts they do not? The mix is a major product decision and it changes the candidate pool entirely.
  2. What are we optimising? The whole section. "Engagement" is not an answer.
  3. What are the non-negotiable constraints? Integrity rules, advertising policy, legal removals, and user controls like muted accounts. These are filters, not preferences.
  4. How many candidates per request, and what is the latency budget?
  5. How is the feed consumed? Infinite scroll, session length, and how often people return — these determine what "a good feed" even means.

The constraints we design against

Invented. The platform is Cobalt, the social network from Sections 5 and 8.

ConstraintValue
Users1.2 billion monthly active
Posts created~500 million per day
Feed requests~300,000 per second at peak
Candidates per request, post-retrieval~1,500
Items per page30, with infinite scroll
Latency budget300 ms p99
Median follows per user~240 accounts

Framing it as machine learning

  • Input: a user, a context, and ~1,500 candidate posts already retrieved.
  • Output: an ordered list of 30 posts.
  • Objective: maximise a weighted combination of predicted outcomes — meaningful interaction, informed consumption, and satisfaction — minus predicted negative outcomes.

Task type: multi-objective ranking over an already-retrieved candidate set.

Three parts of that phrase carry weight.

Multi-objective, because no single predicted outcome describes a good feed. The metrics lesson and Predicting several things at once establish why.

Ranking, because the model's job is ordering, not retrieval. Retrieval was solved by the fan-out design in System Design Interview, plus a recommended-content pool this section treats as an input.

Over an already-retrieved candidate set, which is why this section has no candidate generation lesson and Section 6 does. Section 6 owns two-stage retrieval; this section starts one stage later.

The baseline first

Reverse chronological. Show everything from the accounts a user follows, newest first. No model, no training data, and it has genuine virtues that are worth naming honestly:

  • Fully predictable. A user knows what they will see and can reason about why.
  • No feedback loop, no filter bubble mechanics, no engagement optimisation.
  • Zero cost, and immune to every failure mode in the rest of this section.

Its failures are equally real. A user following 240 accounts may have 900 eligible posts per session and can read 30, so the ordering decides what they see whether or not you call it ranking. Reverse chronological gives that decision to whoever posted most recently, which rewards volume over quality and buries a friend's important news under an account that posts 40 times a day.

Chronological is not "no ranking". It is ranking by recency, and it optimises for posting frequency. Saying that clearly is a much better answer than either dismissing it or romanticising it — and it is worth remembering that offering chronological as a user-selectable option is cheap and is what many platforms do.

Metrics

The central claim of this step: no single metric captures a good feed, and the ones that are easiest to move are the ones most likely to make the product worse.

The ladder, and where it stops being honestClicks and dwell timeComments and resharesMeaningful interactionSurveys and retention
The bottom rungs move fastest under optimisation and are the ones most likely to make the product worse.

The metric ladder

LayerMetricsMoves inRelationship to value
InteractionClick, like, comment, share, replyHoursWeak. Easy to move, easy to game
ConsumptionDwell time, posts read to the end, video watch timeHoursMedium. Correlates with attention, not approval
Meaningful interactionComments and replies between people who know each other, shares with commentary, long-form repliesDaysStronger. Costly signals
Explicit satisfactionSurvey responses, "show more / less of this", savesDaysStrongest per unit, and sparse
Explicit negativesHide, snooze, unfollow, reportDaysVery strong. Deliberate and costly to give
RetentionDaily and monthly return, session frequencyWeeksThe honest long-run measure

Why dwell time and clicks fail

Clicks and likes are cheap actions. Their cost to the user is near zero, so they respond to whatever produces an immediate reaction — which is disproportionately content that provokes. Optimise them and the feed fills with outrage, novelty, and confident claims, because those reliably produce a tap.

Dwell time is better and still fails, in a way worth understanding. Long dwell can mean "this was interesting". It can also mean "this was distressing and I could not look away", or "this was confusing and I read it three times", or "the page was slow". The measurement cannot distinguish them, and a system optimising it will find whichever category is cheapest to produce.

An invented but representative demonstration. Two posts, identical dwell time of 45 seconds:

  • Post A: a long, well-argued piece a user read carefully and then shared with a comment.
  • Post B: an inflammatory claim that made a user angry, which they re-read twice and then hid.

A dwell-time objective rates them identically. Every other signal separates them completely — which is the argument for multi-objective ranking in one example.

Meaningful interaction

The response to that problem is to weight costly signals over cheap ones. A signal is costly when it takes effort or carries social risk:

SignalCost to the userWhat it indicates
A likeNear zeroA reaction occurred
A clickNear zeroCuriosity
A share with no commentLowSome endorsement
A comment of more than a few wordsMediumEngagement with the substance
A reply in an ongoing conversationMedium-highA relationship
A share with commentary to a small groupHighReal endorsement
Responding to a surveyHighConsidered judgement

Weight by cost, not by frequency. This inverts the usual instinct — the rarest signals get the highest weights precisely because they are rare and deliberate.

The failure mode of this approach, which you should name before an interviewer does: optimising comments produces argument, because argument generates comments. A post that makes people angry gets far more replies than one people agree with. So "meaningful interaction" needs its own guard — comment sentiment, whether repliers are connected to each other, and whether conversations end in reports or blocks.

The wellbeing question, as an engineering constraint

Research on social feeds and wellbeing is genuinely mixed and contested; effect sizes vary widely by study, population, and measurement, and any specific figure quoted here would need checking against a current source. What is consistently reported, and what matters for the design, is a divergence: passive consumption of a feed and active interaction with people you know have different relationships to reported wellbeing, and pure-engagement optimisation tends to push a feed toward the former.

The mechanism is not mysterious. Passive consumption is easier to produce, easier to measure, and more responsive to optimisation than the harder-to-manufacture experience of a conversation with someone you know. So an optimiser pointed at engagement drifts toward the passive end.

That produces concrete engineering requirements, and they belong in the design rather than in a closing paragraph:

  1. Distinguish passive from active engagement in the objective. Do not sum them. Weight comments, replies, and shares to known contacts above scroll-past dwell.
  2. Run a continuous satisfaction survey on a sampled slice, and treat it as a release-gate metric with the same standing as engagement.
  3. Make explicit negatives expensive in the score. A hide or an unfollow should outweigh several likes.
  4. Hold out a population on a satisfaction-weighted model as a long-run reference, because two-week experiments cannot measure three-month effects.
  5. Cap the marginal value of session length with a concave transform, so the system does not profit from a user's difficulty stopping.

Data and features

The training set is an impression log produced by the current ranker, which means it describes a feed the system already chose. Everything in this step follows from that.

What the impression log has, and what it lacksPresent in the log• Everything the ranker chose to show• Actions taken on those impressions• Position and viewport dwellAbsent from the log• Posts the ranker never surfaced• Whether a skip meant dislike or scroll• Value that produced no visible action
A negative has to be defined, not found — shown and scrolled past is a label, never shown is not.

The features that matter

Author–viewer affinity is the strongest feature family in a feed, more predictive than content features by a wide margin:

  • Interaction history: how often this viewer has liked, commented on, or shared this author's posts, over several trailing windows.
  • Interaction recency: days since the last interaction with this author.
  • Symmetry: does the author engage back? Mutual interaction is a much stronger signal than one-directional.
  • Relationship type: mutual follow, one-way follow, or no follow at all (for recommended content).
  • Profile-visit history: viewing someone's profile is a strong, deliberate interest signal.
  • Shared context: mutual connections, shared groups, co-location.

A practical note worth stating: affinity is expensive to compute per (viewer, author) pair and there are billions of pairs, so it is precomputed for the author set a user actually follows and computed live only for the recommended pool.

Content features:

  • Type — text, image, video, link, poll. Each has different baseline engagement, so the model must be able to condition on it.
  • Text and image embeddings, precomputed at post time.
  • Topic classification.
  • Quality signals: originality (is this a reposted image), clickbait score, integrity model scores from Section 5, whether the link domain is reputable.
  • Early engagement velocity, from the first minutes after posting — powerful, and the source of a rich-get-richer loop.

Recency. A feed post's value decays fast, and the decay rate is content-dependent: a news link is worthless in 12 hours, a friend's life announcement is relevant for days. Encode age bucketed rather than raw, and cross it with content type.

Context. Time of day, device, session position (the 3rd item and the 60th are different problems), connection quality, and what the user has already seen this session.

Building a training set from a biased log

Every example in the log is an impression the current ranker chose to show. Three specific biases, each with a correction:

Position bias. As in Sections 4, 6 and 8. Include position as a training feature, fix it to a reference value at serving time, or reweight by inverse propensity.

Selection bias — the severe one here. A user with 900 eligible posts sees 30. The 870 they never saw produce no data at all. The model is trained exclusively on the current ranker's choices, so it learns to reproduce them. This is the mechanism by which a feed ranker becomes progressively more confident about a progressively narrower slice of content.

The mitigation is a randomised holdout: for a small percentage of requests, insert a few randomly-selected eligible posts at random positions. The engagement on those is unbiased and it is the only data that tells you about content the ranker would never have picked. It costs measurable short-term engagement and it is not optional — without it, the model has no way to discover it is wrong.

Exposure duration bias. A post at the top of the feed is on screen while the user decides what to do; one further down may be scrolled past in 200 milliseconds. "Not engaged with" means different things at different positions. Log viewport time per impression and use it to weight negatives: a post that was on screen for 3 seconds and not engaged with is a real negative; one visible for 200 milliseconds is not.

Negatives, precisely

SignalLabel
Engaged (any positive action)Positive, graded by the action's cost
On screen ≥ 2 s, no actionNegative
On screen < 500 msDrop — not evidence
Hidden, snoozed, unfollowed, reportedStrong negative, weighted heavily
Never renderedNot in the training set at all

The third and fifth rows are the ones people get wrong, and they matter more here than in most problems because feed scrolling is fast and the majority of impressions are brief.