Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
News feed: multi-objective heads and the shared-trunk model
The mechanism at the centre of this section: predict each outcome separately, then combine. The first half of this lesson builds the heads and the scoring formula, and shows where the platform's values actually live — in a handful of weights.
The second half is the model underneath: eight heads, one shared representation, and the specific way that arrangement goes wrong.
Why not one model predicting "value"
The tempting design is a single model predicting one number — how good this post is for this user. It fails for three reasons.
There is no label for it. Nobody records "value". Every label available is a specific action, so a single-output model has to be trained on some particular action, which makes it a model of that action wearing a general name.
The trade-offs become invisible. With one number, the decision to weight comments above likes is buried inside learned weights where nobody can see, discuss, or change it. With separate heads and explicit weights, that decision is a line of configuration a product team can argue about and adjust in an afternoon.
Different outcomes need different data. Reports are 10,000 times rarer than likes. A single model optimising a blended label learns almost nothing about reports. Separate heads let each task use its own loss weighting, its own sampling, and its own evaluation.
The heads
For Cobalt's feed, eight heads over a shared representation:
| Head | Predicts | Base rate | Role |
|---|---|---|---|
| P(click) | Opens the post or its link | ~4% | Weak positive |
| P(dwell > 5 s) | Meaningful reading time | ~11% | Consumption |
| P(like) | Any reaction | ~2.5% | Weak positive |
| P(comment) | Writes a comment | ~0.3% | Costly positive |
| P(share) | Shares, especially with commentary | ~0.15% | Costly positive |
| P(hide) | Hides or snoozes | ~0.08% | Strong negative |
| P(report) | Reports the post | ~0.004% | Very strong negative |
| P(unfollow) | Unfollows the author after seeing it | ~0.01% | Very strong negative |
Note the range: the most common head is roughly a thousand times more frequent than the rarest. That imbalance is the central training difficulty and The model lesson addresses it.
Combining into one score
score = w_click·P(click) + w_dwell·P(dwell>5s) + w_like·P(like) + w_comment·P(comment) + w_share·P(share) − w_hide·P(hide) − w_report·P(report) − w_unfollow·P(unfollow)An illustrative weight set, invented, to make the shape concrete. The last column is each positive head's share of the positive part of a typical score:
| Head | Weight | Weight × base rate | Share of typical score |
|---|---|---|---|
| click | 1 | 0.040 | 8% |
| dwell > 5 s | 2 | 0.220 | 44% |
| like | 3 | 0.075 | 15% |
| comment | 30 | 0.090 | 18% |
| share | 50 | 0.075 | 15% |
| hide | −200 | −0.160 | Negative |
| report | −2,000 | −0.080 | Negative |
| unfollow | −1,000 | −0.100 | Negative |
Three things to read from that table, and all three are worth saying aloud.
Rare signals need large weights to matter at all. A comment weight of 30 against a like weight of 3 does not mean a comment is worth ten likes in importance — it means that after multiplying by base rates, comments contribute a comparable share of the score. Look at the third column, not the second, when reasoning about what the model is actually optimising.
Negative weights are enormous. A single predicted report at −2,000 outweighs everything positive a post could offer. That is deliberate: showing content someone reports is a failure no amount of engagement compensates for.
The weights are product policy. No dataset says a share is worth 50 units of click. Someone decides. That decision is where the platform's values are actually encoded, far more than in any published policy document, and being able to say so is a mark of seniority.
Setting and tuning the weights
Four approaches, and a recommendation:
- Hand-set from product principles. Start here. Write down the ordering you want — costly signals above cheap ones, negatives dominant — and pick numbers that produce that ordering after base-rate multiplication.
- Tune by A/B test. Run several weight sets and compare on the metric you actually care about — retention, satisfaction — not on the components. This is how weights are set in practice. It is slow, since each test needs weeks, so the search must be coarse.
- Fit to a long-term objective. Train a model predicting 28-day retention from a session's composition, then choose the weights that maximise it. Principled and it needs a great deal of data and a long feedback loop.
- Per-user weights. Learn that some users respond to different things. Powerful, and it risks giving each user more of what they already do, which is the filter bubble mechanism with extra machinery.
Recommendation: hand-set from principles, tune coarsely with A/B tests on retention and satisfaction, and hold the negative weights fixed at high values rather than letting a tuning process erode them. That last constraint matters — an optimiser searching for short-term engagement will always want to reduce the penalty on hide and report.
Calibration matters here
The score is a weighted sum of probabilities, which means the heads' values are being combined arithmetically. If P(comment) is systematically inflated by 30% relative to P(like), the effective weights are not what the configuration says.
Check calibration per head after every retrain and apply a per-head recalibration layer if needed. The ad click metrics lesson covers the diagnostics; the point here is that multi-objective scoring makes calibration a correctness requirement rather than an optional nicety.
The model: why share a trunk
Eight heads need a model underneath them. The alternative to sharing is eight independent models. It works, and it is worse for three reasons:
Cost. Eight models over 1,500 candidates is eight forward passes. A shared trunk with eight small heads is one pass and eight cheap projections — close to an eight-fold saving on the dominant cost in the request path.
Data efficiency. P(report) has a 0.004% base rate. Trained alone, it sees very few positives and cannot learn a good representation of a post from them. Sharing a trunk with P(click), which has a thousand times more data, gives it a representation learned from all of it.
Consistency. Eight independent models can disagree in incoherent ways — one confident the post is great, another confident it will be reported, with no shared understanding of the post underneath. A shared representation makes the heads disagreements about outcomes rather than about the input.
The architecture
- Inputs: viewer features, author features, affinity features, content features and embeddings, and context features. High-cardinality identifiers pass through embedding tables (see ad click prediction: features).
- Shared trunk: several dense layers producing a representation of this (viewer, post) pair.
- Per-task towers: one or two small layers per objective, so each head has capacity that does not compete with the others.
- Heads: eight sigmoid outputs.
The per-task towers matter more than they look. Without them, every task-specific pattern has to fit inside the shared trunk, and the tasks fight over its capacity. With them, the trunk learns what all tasks need and each tower learns what only its own task needs.
Negative transfer
Four mitigations, in the order to try them:
- Weight the per-task losses so no task dominates. Inverse base rate is a starting point, then tune. This is the first and often sufficient fix.
- Give each head its own tower, as above.
- Use gated experts. Instead of one shared trunk, have several parallel expert sub-networks plus a small per-task gate that learns how much of each expert its task should use. Tasks that need different features can lean on different experts while still sharing where sharing helps. This is the standard architectural response to negative transfer in ranking models, and describing the mechanism — parallel experts, per-task gating — is better than naming a specific published model.
- Split the tasks. Nothing requires all eight in one model. If P(report) and P(hide) genuinely conflict with the engagement heads, a separate integrity model is a legitimate answer — and one that connects cleanly to Section 5, which already builds one.
Training
- Loss: per-head binary cross-entropy, summed with task weights.
- Sampling: shared batches, so every example contributes to every head. Rare positives get their weight from the loss weighting rather than from resampling, which keeps the batches consistent across heads.
- Split by time, with users appearing in only one side of the split.
- Evaluate per head, against both single-task baselines and the previous production model. An aggregate score across eight heads hides exactly the regressions that matter.