Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
News feed: framing, costly signals and a biased impression log
Prompt: "Design the ranking system for a personalised news feed."
This section owns multi-objective optimisation — predicting several outcomes at once and combining them into a single score. It is the technique Section 6 used and deferred, and it is the heart of how every large feed is actually ranked.
This first lesson sets up why a single objective will not do. It frames the problem, makes the case that no single metric captures a good feed, and builds the training data from an impression log that describes a feed the system already chose.
The relationship to the news feed design in System Design Interview
This section and the news feed case study in System Design Interview design the same product and solve opposite halves of it. Saying so out loud in an interview is a genuinely good move, because it shows you know which problem you are being asked about.
| System Design Interview: news feed | This section | |
|---|---|---|
| The question | How does a post get from an author to a follower's feed? | Given the posts that could be shown, in what order? |
| The core trade-off | Fan-out on write versus fan-out on read | Which objectives, and with what weights |
| The hard part | Celebrity accounts with 40 million followers | Objectives that pull against each other |
| The output | A set of candidate post IDs, unordered | An ordered list of 30 |
| Failure mode | Feed loads slowly, or misses posts | Feed loads fast and is not worth reading |
The system design version produces the candidate set this section ranks. A candidate who has done both can say: "the fan-out design gives me a few hundred to a couple of thousand candidates per request; my job here starts from there." That sentence connects two rounds of preparation and takes four seconds.
The clarifying questions
- Which content sources? Only posts from accounts the user follows, or also recommended content from accounts they do not? The mix is a major product decision and it changes the candidate pool entirely.
- What are we optimising? The whole section. "Engagement" is not an answer.
- What are the non-negotiable constraints? Integrity rules, advertising policy, legal removals, and user controls like muted accounts. These are filters, not preferences.
- How many candidates per request, and what is the latency budget?
- How is the feed consumed? Infinite scroll, session length, and how often people return — these determine what "a good feed" even means.
The constraints we design against
Invented. The platform is Cobalt, the social network from Sections 5 and 8.
| Constraint | Value |
|---|---|
| Users | 1.2 billion monthly active |
| Posts created | ~500 million per day |
| Feed requests | ~300,000 per second at peak |
| Candidates per request, post-retrieval | ~1,500 |
| Items per page | 30, with infinite scroll |
| Latency budget | 300 ms p99 |
| Median follows per user | ~240 accounts |
Framing it as machine learning
- Input: a user, a context, and ~1,500 candidate posts already retrieved.
- Output: an ordered list of 30 posts.
- Objective: maximise a weighted combination of predicted outcomes — meaningful interaction, informed consumption, and satisfaction — minus predicted negative outcomes.
Task type: multi-objective ranking over an already-retrieved candidate set.
Three parts of that phrase carry weight.
Multi-objective, because no single predicted outcome describes a good feed. The metrics lesson and Predicting several things at once establish why.
Ranking, because the model's job is ordering, not retrieval. Retrieval was solved by the fan-out design in System Design Interview, plus a recommended-content pool this section treats as an input.
Over an already-retrieved candidate set, which is why this section has no candidate generation lesson and Section 6 does. Section 6 owns two-stage retrieval; this section starts one stage later.
The baseline first
Reverse chronological. Show everything from the accounts a user follows, newest first. No model, no training data, and it has genuine virtues that are worth naming honestly:
- Fully predictable. A user knows what they will see and can reason about why.
- No feedback loop, no filter bubble mechanics, no engagement optimisation.
- Zero cost, and immune to every failure mode in the rest of this section.
Its failures are equally real. A user following 240 accounts may have 900 eligible posts per session and can read 30, so the ordering decides what they see whether or not you call it ranking. Reverse chronological gives that decision to whoever posted most recently, which rewards volume over quality and buries a friend's important news under an account that posts 40 times a day.
Chronological is not "no ranking". It is ranking by recency, and it optimises for posting frequency. Saying that clearly is a much better answer than either dismissing it or romanticising it — and it is worth remembering that offering chronological as a user-selectable option is cheap and is what many platforms do.
Metrics
The central claim of this step: no single metric captures a good feed, and the ones that are easiest to move are the ones most likely to make the product worse.
The metric ladder
| Layer | Metrics | Moves in | Relationship to value |
|---|---|---|---|
| Interaction | Click, like, comment, share, reply | Hours | Weak. Easy to move, easy to game |
| Consumption | Dwell time, posts read to the end, video watch time | Hours | Medium. Correlates with attention, not approval |
| Meaningful interaction | Comments and replies between people who know each other, shares with commentary, long-form replies | Days | Stronger. Costly signals |
| Explicit satisfaction | Survey responses, "show more / less of this", saves | Days | Strongest per unit, and sparse |
| Explicit negatives | Hide, snooze, unfollow, report | Days | Very strong. Deliberate and costly to give |
| Retention | Daily and monthly return, session frequency | Weeks | The honest long-run measure |
Why dwell time and clicks fail
Clicks and likes are cheap actions. Their cost to the user is near zero, so they respond to whatever produces an immediate reaction — which is disproportionately content that provokes. Optimise them and the feed fills with outrage, novelty, and confident claims, because those reliably produce a tap.
Dwell time is better and still fails, in a way worth understanding. Long dwell can mean "this was interesting". It can also mean "this was distressing and I could not look away", or "this was confusing and I read it three times", or "the page was slow". The measurement cannot distinguish them, and a system optimising it will find whichever category is cheapest to produce.
An invented but representative demonstration. Two posts, identical dwell time of 45 seconds:
- Post A: a long, well-argued piece a user read carefully and then shared with a comment.
- Post B: an inflammatory claim that made a user angry, which they re-read twice and then hid.
A dwell-time objective rates them identically. Every other signal separates them completely — which is the argument for multi-objective ranking in one example.
Meaningful interaction
The response to that problem is to weight costly signals over cheap ones. A signal is costly when it takes effort or carries social risk:
| Signal | Cost to the user | What it indicates |
|---|---|---|
| A like | Near zero | A reaction occurred |
| A click | Near zero | Curiosity |
| A share with no comment | Low | Some endorsement |
| A comment of more than a few words | Medium | Engagement with the substance |
| A reply in an ongoing conversation | Medium-high | A relationship |
| A share with commentary to a small group | High | Real endorsement |
| Responding to a survey | High | Considered judgement |
Weight by cost, not by frequency. This inverts the usual instinct — the rarest signals get the highest weights precisely because they are rare and deliberate.
The failure mode of this approach, which you should name before an interviewer does: optimising comments produces argument, because argument generates comments. A post that makes people angry gets far more replies than one people agree with. So "meaningful interaction" needs its own guard — comment sentiment, whether repliers are connected to each other, and whether conversations end in reports or blocks.
The wellbeing question, as an engineering constraint
Research on social feeds and wellbeing is genuinely mixed and contested; effect sizes vary widely by study, population, and measurement, and any specific figure quoted here would need checking against a current source. What is consistently reported, and what matters for the design, is a divergence: passive consumption of a feed and active interaction with people you know have different relationships to reported wellbeing, and pure-engagement optimisation tends to push a feed toward the former.
The mechanism is not mysterious. Passive consumption is easier to produce, easier to measure, and more responsive to optimisation than the harder-to-manufacture experience of a conversation with someone you know. So an optimiser pointed at engagement drifts toward the passive end.
That produces concrete engineering requirements, and they belong in the design rather than in a closing paragraph:
- Distinguish passive from active engagement in the objective. Do not sum them. Weight comments, replies, and shares to known contacts above scroll-past dwell.
- Run a continuous satisfaction survey on a sampled slice, and treat it as a release-gate metric with the same standing as engagement.
- Make explicit negatives expensive in the score. A hide or an unfollow should outweigh several likes.
- Hold out a population on a satisfaction-weighted model as a long-run reference, because two-week experiments cannot measure three-month effects.
- Cap the marginal value of session length with a concave transform, so the system does not profit from a user's difficulty stopping.
Data and features
The training set is an impression log produced by the current ranker, which means it describes a feed the system already chose. Everything in this step follows from that.
The features that matter
Author–viewer affinity is the strongest feature family in a feed, more predictive than content features by a wide margin:
- Interaction history: how often this viewer has liked, commented on, or shared this author's posts, over several trailing windows.
- Interaction recency: days since the last interaction with this author.
- Symmetry: does the author engage back? Mutual interaction is a much stronger signal than one-directional.
- Relationship type: mutual follow, one-way follow, or no follow at all (for recommended content).
- Profile-visit history: viewing someone's profile is a strong, deliberate interest signal.
- Shared context: mutual connections, shared groups, co-location.
A practical note worth stating: affinity is expensive to compute per (viewer, author) pair and there are billions of pairs, so it is precomputed for the author set a user actually follows and computed live only for the recommended pool.
Content features:
- Type — text, image, video, link, poll. Each has different baseline engagement, so the model must be able to condition on it.
- Text and image embeddings, precomputed at post time.
- Topic classification.
- Quality signals: originality (is this a reposted image), clickbait score, integrity model scores from Section 5, whether the link domain is reputable.
- Early engagement velocity, from the first minutes after posting — powerful, and the source of a rich-get-richer loop.
Recency. A feed post's value decays fast, and the decay rate is content-dependent: a news link is worthless in 12 hours, a friend's life announcement is relevant for days. Encode age bucketed rather than raw, and cross it with content type.
Context. Time of day, device, session position (the 3rd item and the 60th are different problems), connection quality, and what the user has already seen this session.
Building a training set from a biased log
Every example in the log is an impression the current ranker chose to show. Three specific biases, each with a correction:
Position bias. As in Sections 4, 6 and 8. Include position as a training feature, fix it to a reference value at serving time, or reweight by inverse propensity.
Selection bias — the severe one here. A user with 900 eligible posts sees 30. The 870 they never saw produce no data at all. The model is trained exclusively on the current ranker's choices, so it learns to reproduce them. This is the mechanism by which a feed ranker becomes progressively more confident about a progressively narrower slice of content.
The mitigation is a randomised holdout: for a small percentage of requests, insert a few randomly-selected eligible posts at random positions. The engagement on those is unbiased and it is the only data that tells you about content the ranker would never have picked. It costs measurable short-term engagement and it is not optional — without it, the model has no way to discover it is wrong.
Exposure duration bias. A post at the top of the feed is on screen while the user decides what to do; one further down may be scrolled past in 200 milliseconds. "Not engaged with" means different things at different positions. Log viewport time per impression and use it to weight negatives: a post that was on screen for 3 seconds and not engaged with is a real negative; one visible for 200 milliseconds is not.
Negatives, precisely
| Signal | Label |
|---|---|
| Engaged (any positive action) | Positive, graded by the action's cost |
| On screen ≥ 2 s, no action | Negative |
| On screen < 500 ms | Drop — not evidence |
| Hidden, snoozed, unfollowed, reported | Strong negative, weighted heavily |
| Never rendered | Not in the training set at all |
The third and fifth rows are the ones people get wrong, and they matter more here than in most problems because feed scrolling is fast and the majority of impressions are brief.