Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Harmful content: tiered serving, adversaries and appeals
500 million posts a day and one multimodal model that costs real money per inference. The arithmetic decides the architecture, and the first half of this lesson works it through.
The second half deals with the fact that makes this system's monitoring different from every other case study in the course: it has an adversary. People actively work out what the classifier catches and change their content to avoid it, while appeals, regional law and fairness audits keep the system accountable.
The cost that forces the shape
Take the hybrid fusion model from the model lesson at an illustrative £0.0008 per post — a few hundred milliseconds of accelerator time for a post with an image and a short video.
Run it on every post: 500,000,000 × £0.0008 = £400,000 per day, roughly £146 million a year, for a model that will be replaced in eighteen months.
That is not affordable, and it is also unnecessary: the overwhelming majority of posts are a photo of a meal or a comment about football. The expensive model's capability is wasted on them.
The tiers
| Tier | What it is | Runs on | Unit cost | Daily cost | Passes on |
|---|---|---|---|---|---|
| 0 | Hash match + keyword rules + author reputation | 500M (all) | ~£0.0000002 | £100 | ~480M |
| 1 | Cheap model: small text encoder + low-res image encoder, single score per category | 480M | ~£0.00002 | £9,600 | ~15M (3%) |
| 2 | Full hybrid multimodal model | 15M | ~£0.0008 | £12,000 | ~750k (0.15%) |
| 3 | Human review, prioritised | 750k | ~£0.05 | £37,500 | — |
| Total | ~£59,200/day |
Against £400,000 a day for the naive design, the tiered pipeline costs about 15% — and the human tier, at £37,500, is the largest line. That ordering recurs across this course: at maturity, people cost more than accelerators.
All figures are invented and chosen to show the shape of the calculation. The shape is what transfers.
How the tiers work together
Tier 0 is instantaneous and runs synchronously before publication. A hash match against known-violating content blocks the post outright. This is the only tier that can act pre-publication without adding user-visible latency.
Tier 1 runs asynchronously within a couple of seconds. Its job is recall, not precision — it must not discard anything the heavy model would have caught. Tune it at a very low threshold: passing 3% of traffic to tier 2 is a design parameter, and you would rather pass 5% than miss. Measure tier 1 by recall relative to tier 2's decisions on a sample where both ran, which is a specific and unusual metric worth naming.
Tier 2 produces the calibrated per-category scores that drive the threshold table from the metrics lesson. Auto-remove above the high threshold, demote in the middle band, escalate the uncertain band.
Tier 3 is human review, and its queue ordering is where the remaining leverage lives.
Prioritising the human queue
A million reviews a day against a queue that could be many times larger. Order by expected harm prevented, not by score and not by arrival time:
priority = P(violating) × severity(category) × expected_reach × time_decay- P(violating) — the tier 2 score.
- severity — the per-category weight from the cost matrix.
- expected reach — predicted views if left up, from author follower count, early engagement velocity, and whether the post is being amplified by recommendations. A post already accelerating deserves review before an identical post that is not.
- time decay — harm accumulates while a post is up, so a post's priority rises with age in the queue, which also prevents starvation.
The reach term is the one candidates miss and it is the most valuable. A violating post with 0.6 confidence going viral does far more harm than one with 0.95 confidence seen by nine people.
Also required: routing by specialisation. Child safety and self-harm go to trained specialist queues with support provisions, not to the general pool. Language routing matters too — a reviewer who does not read the language cannot judge context.
Adversarial evasion
With the pipeline running, the monitoring problem begins — and here it has an opponent.
Everything else in this course degrades passively as the world changes. Here, people actively work out what the classifier catches and change their content to avoid it.
The techniques are well known and cheap:
- Text perturbation. Character substitution, deliberate misspelling, inserted punctuation, homoglyphs from other alphabets that render identically.
- Image perturbation. Small crops, colour shifts, added borders, overlaid noise — enough to defeat hash matching while leaving the content recognisable to a person.
- Text as image. Rendering a message as a picture to bypass text classifiers entirely.
- Context splitting. Benign post, harmful comment. Benign image, harmful caption in a reply.
- Coded language. New euphemisms that carry the meaning to an in-group and nothing to a classifier. These evolve continuously.
The consequences for the design:
- Retraining cadence is set by the adversary, not by data volume. Weekly at minimum for the fastest-moving categories.
- Augment training with perturbations. Train on character-substituted and visually-perturbed versions of known positives so the model learns invariance to the cheapest evasions.
- Do not publish thresholds or scores. Any feedback about how close a post came to being flagged is a gradient signal handed to an adversary.
- Rate-limit the implicit oracle. An account that posts fifty near-identical variants is probing the classifier. That pattern is itself a detection signal.
- Rely on non-content signals. Account age, posting velocity, network structure, and device fingerprints are far harder to perturb than pixels and characters.
Distribution shift as new harm types emerge
New harms appear that no category covers: a new scam format, a new coordinated campaign, a new product being sold illegally. The model has never seen them and its scores on them are near zero, so no monitor keyed to model output will fire.
Detection has to come from outside the model:
- User report volume by topic. Reports are noisy for individual decisions and excellent as an aggregate early warning. A spike in reports on posts the model scores low is the clearest signal a new harm type exists.
- The random review sample. Reviewers see things the model does not suspect, and their "violating, but no category fits" responses are the leading indicator.
- Clustering of low-score reported content. Group reported-but-not-flagged posts by embedding and look for dense clusters. A new campaign appears as a cluster.
Adding a category is then a bounded project: policy definition, a labelling push, a new head on the existing trunk, threshold-setting, and a specialised queue.
Appeals
Appeals are both a fairness requirement and the best labelled data in the system.
Every appeal produces a second, independent judgement on a case where the system acted. The overturn rate is your honest false positive measure (see the metrics lesson). Overturned cases are high-value training examples: they are precisely the boundary cases the model gets wrong.
Two design requirements. Route appeals to a different reviewer than the original decision, or the second judgement is not independent. And track overturn rate by category, language, and region — a spike in one segment means the model is systematically over-enforcing against one group, which is exactly the failure that turns a moderation system into a discrimination mechanism.
Per-region policy differences
What is unlawful differs by country. Some speech is protected in one jurisdiction and prohibited in another. Local law may require removal within a stated window, and may require retention of removed content for investigation.
The architecture that handles this: one global model producing scores, per-region policy configuration deciding actions. Regional thresholds, regional category enablement, regional retention rules — all configuration, not model weights. Retraining a model per country is untenable at forty countries; reconfiguring thresholds is a deployment.
Where a region genuinely needs different detection rather than different enforcement — a category that exists only there — add a head, not a model.
Auditing for demographic bias
This is a design requirement, and it belongs here rather than in a closing caveat.
Text classifiers trained on moderation labels have repeatedly been found to over-flag content in some dialects and from some communities — including, in published research on toxicity classifiers, content written in African-American English and content in which slurs are reclaimed by the groups they target. Specific figures vary by study, dataset, and model, and any number would need checking against a current source; the direction of the finding is consistent enough to design against.
The mechanism is understandable: annotators unfamiliar with a dialect or a community's norms label its ordinary speech as aggressive, and the model learns that.
What the design must include:
- Stratified evaluation sets by language, dialect, and region, with per-slice precision and recall reported at every release.
- Appeal overturn rate segmented the same way, monitored continuously — it is the production version of the same measurement.
- Annotator pools that include speakers of the dialects being judged, and policy guidance that explicitly addresses reclamation and counter-speech.
- A release gate: no slice may exceed the global false positive rate by more than a stated margin.
- A published appeals path, because no amount of measurement removes every error and the people affected need recourse.
Over-enforcement against a community is not a side effect to be noted. It is the system failing, in a way that removes those users' ability to use the platform.