Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Harmful content: framing, per-category thresholds and labels


Prompt: "Design a system that detects harmful content on a social platform."

This is the most operationally realistic problem in the course. Both error types hurt real people, the positive class is vanishingly rare, and the model is one component of a system that includes humans, appeals, and per-country policy.

This lesson covers the first three steps, and in this problem they carry more weight than usual: what happens to a positive decides the precision you need, the cost of each error differs per harm category, and the labels come from a moderation queue with all the selection effects that implies.

From upload to enforcementContent uploadedHash match and filtersMultimodal classifierPer-category thresholdsAuto-remove or human queue
Framing it as multi-label rather than one harmful score is what lets each category carry its own threshold and its own policy.

The clarifying questions

  1. Which harm categories? "Harmful" is not one thing. Graphic violence, hate speech, harassment, adult nudity, self-harm promotion, scams, regulated goods, and coordinated misinformation are eight different problems with different base rates, different label sources, and different acceptable error rates. Ask which ones are in scope.
  2. Which modalities? Text, images, video, audio, and — critically — the combination. A benign image with benign text can be harmful together. If that case is in scope, the architecture changes fundamentally.
  3. What happens to a positive? This is the highest-leverage question in the list. Auto-remove is irreversible and demands high precision. Demote or age-gate is reversible and tolerates much lower precision. Route to human review costs money but tolerates the lowest precision of all. Most real systems do all three, at different confidence bands.
  4. What is the volume, and how fast must detection be? Before publication, or within minutes after?
  5. Is there an appeals process? If yes, false positives are recoverable and the threshold can be more aggressive. If no, every false positive is permanent.

The constraints we design against

Invented. The platform is Cobalt, a social network with text, image, and video posts.

ConstraintValue
Volume~500 million posts per day
ModalitiesText, image, video (average 40 s), and combinations
Base rate of violating content~0.2% overall; individual categories from 0.001% to 0.08%
Detection latencyWithin 60 s of posting for the automated path
Categories in scopeEight, with separate policies
Human moderation capacity~1 million reviews per day
AppealsYes, with a stated 48-hour turnaround

Framing it as machine learning

  • Input: a post — text, zero or more images, an optional video, plus author and context features.
  • Output: for each of the eight harm categories, a probability that the post violates that policy.
  • Objective: rank posts by the expected harm of leaving them up, so that limited automated and human enforcement capacity is spent where it does most good.

The task type is multi-label classification with a human in the loop.

Three things in that phrase, each doing work.

Multi-label, not multi-class. A post can violate several policies at once — a harassing message containing a scam link. Multi-class forces one label and would discard information. Eight independent binary outputs is the right shape.

With a human in the loop. This is not a decoration on the framing; it changes the metric discussion entirely. The model's output is not a verdict. It is a triage decision among four actions: allow, demote, remove, or escalate to a person. A model with 40% precision is useless as an auto-remover and excellent as a queue prioritiser, and those are the same model at different thresholds.

Ranking harm, not detecting it. Because enforcement capacity is finite — 1 million reviews a day against 500 million posts — the real job is allocation. Framing it as allocation rather than detection leads to better answers about thresholds, prioritisation, and cost.

The baseline first

Three layers, none of them a neural network, and together they catch a substantial share of the real volume:

  1. Hash matching. Content identical or near-identical to something already actioned. Uses perceptual hashing so small edits — a crop, a re-encode, added text — still match. This is the highest-precision signal in the entire system and it costs microseconds. A large share of violating uploads are re-uploads of known material.
  2. Keyword and regex matching on text, plus known-bad link domains. High precision on explicit slurs and known scam infrastructure, easily evaded by spelling variants, and prone to catching quotation and counter-speech.
  3. Author reputation. Accounts with prior violations, very new accounts, and accounts with abnormal posting rates. Not content-based at all, and a strong prior.

These three run first in the pipeline and stay there permanently. They are not scaffolding to be replaced. Presenting them as a baseline that a model will supersede is a mistake — the model handles what these cannot, and they handle what the model should not waste capacity on.

Metrics under a real cost asymmetry

Section 3 (Street View blurring) had an asymmetry so extreme that recall dominated everything. Here both errors are genuinely expensive, in different currencies, and the balance changes per category.

One threshold per category, not one for the systemLowExtreme0.15LowSevere0.20HighHigh0.60HighLow0.85FP costFN costThresholdChild safetyTerrorismHate speechSpamThreshold moves with the ratio of the two costs, not with accuracy.
Both errors are expensive here, in different currencies, so a single global threshold is wrong for every category at once.

What each error costs

A false negative leaves harmful content visible. The cost scales with reach — a violating post seen by 12 people and one seen by 400,000 are not comparable — and with severity. Content promoting self-harm that reaches a vulnerable person is not on the same scale as a spam comment.

A false positive removes or suppresses legitimate speech. It silences a person who did nothing wrong, generates an appeal, and — repeated at scale on a particular community, dialect, or topic — becomes a systematic bias with real consequences for who can use the platform.

Neither is a rounding error. This is the honest version, and treating either as a nuisance is how moderation systems go wrong.

Why a single global threshold is wrong

Suppose one threshold of 0.85 across all eight categories. Consider two of them.

Child safety content. Severity is extreme, volume is low, and the cost of a false negative is unbounded. The right threshold is very low — escalate anything remotely suspicious to a specialist human queue. A false positive costs a few minutes of review.

Spam. Severity is low, volume is enormous, and both errors are cheap. Automated action at moderate confidence is fine, and human review is not worth its cost at all.

One threshold cannot serve both. At 0.85, child-safety content is under-enforced and spam is over-reviewed. The thresholds are not a hyperparameter to tune globally; they are per-category policy decisions, and the argument for that is a strong thing to make in an interview.

Setting thresholds from a cost matrix

Give each category four numbers — the cost of a false negative, the cost of a false positive, the base rate, and the daily volume — and derive the operating points.

CategoryFN costFP costBase rateAuto-remove aboveHuman review bandDemote band
Child safetyExtremeLow0.001%Never auto — always human> 0.10—
Graphic violenceHighMedium0.02%0.970.55 – 0.970.30 – 0.55
Hate speechHighHigh0.04%0.980.45 – 0.980.25 – 0.45
HarassmentHighHigh0.05%Never auto0.40 – 1.000.20 – 0.40
Adult nudityMediumMedium0.08%0.950.70 – 0.950.40 – 0.70
Self-harmExtremeMedium0.006%Never auto> 0.15—
ScamsMediumLow0.06%0.900.60 – 0.900.35 – 0.60
MisinformationMediumVery high0.03%Never auto0.60 – 1.000.35 – 0.60

All figures invented. Three patterns in the table are the actual content:

  • Categories where a false positive is a speech harm — hate speech, harassment, misinformation — either never auto-remove or require near-certainty. Uncertainty routes to a human or to demotion, both reversible.
  • Categories where a false negative is catastrophic — child safety, self-harm — have very low review thresholds and accept a high review volume as the price.
  • Low-severity, high-volume categories automate aggressively, because human review does not repay its cost there.

The metrics that ship

MetricPer categoryWhy
Recall at fixed precisionYes"Recall at 95% precision" is directly actionable for an auto-remove threshold
Precision at fixed recallYesUsed where recall is mandated by policy
PR-AUCYesThe threshold-free comparison across models. Never ROC-AUC — see Step 2: offline metrics
PrevalenceYesEstimated fraction of views that land on violating content — the true north star
Appeal overturn rateYesThe share of enforcement decisions reversed on appeal. The honest false positive measure
Review queue volume and waitOverallOperational; determines whether the thresholds are affordable

Prevalence deserves emphasis. Recall measures what fraction of violating posts you caught, weighting a post seen by 12 people the same as one seen by 400,000. Prevalence — measured by sampling actual views and having them reviewed — measures what fraction of the user experience is harmful. That is the thing the product cares about, and a system can improve recall while prevalence worsens if it is catching low-reach content and missing viral content.

Appeal overturn rate is the false positive measure you can actually get. Precision on a labelled set tells you about your labelled set; overturn rate tells you how often you were wrong about a real person's real post, as judged by a second reviewer. Track it per category and per user segment.

Data and labels

The label source here is a human moderation queue, which means the training data is a record of what moderators were shown and how they judged it — with all the selection effects that implies.

What the moderation queue actually recordsAll uploadsFlagged byusers or modelSampledinto the queueModeratorjudgementTraining labelNothing the old model ignored ever reaches a moderator.
The label set is a record of what moderators were shown, so it inherits every selection effect of the system that fed them.

Where labels come from

Enforcement decisions. Every human review produces a labelled example: this post, this category, violating or not. Volume is large — around 1 million a day at Cobalt — and quality is high, since reviewers are trained against a written policy.

The selection bias is severe and worth stating plainly: reviewers only see posts that were routed to them, and routing is done by the current model and by user reports. Content the model does not suspect and nobody reports never enters the label set. So the labelled data systematically over-represents what the current system already finds.

User reports. A user flags a post. High volume, very low precision — most reports are disagreement, dislike, or coordinated brigading rather than policy violation. Useful as a routing signal, dangerous as a training label without human confirmation.

Deliberate random sampling. A random sample of all posts — say 20,000 a day — sent for review regardless of model score. Low yield: at a 0.2% base rate, 20,000 posts contain about 40 violations. Expensive per positive, and the only unbiased data in the system. It is what prevalence estimation requires and what tells you about the content the model never suspects. Never cut it.

Weak supervision. Posts from accounts later banned for a category; posts matching known-bad hashes; content removed by legal request. Noisy, plentiful, useful for bootstrapping a new category.

Inter-annotator disagreement

Some content is genuinely ambiguous, and the disagreement is information rather than noise.

Reviewers agree strongly on unambiguous categories — explicit nudity, graphic violence — and much less on judgement-heavy ones. Hate speech depends on who is speaking, about whom, in what context, and whether it is a quotation, a reclamation of a slur, or counter-speech. Harassment depends on a history between two accounts that a reviewer looking at one post cannot see.

What to do about it:

  • Multiple reviewers on ambiguous categories, with the label being the majority and the disagreement recorded. Costs 2–3× per label, and buys a usable ceiling.
  • Train on soft labels. If two of three reviewers said "violating", the target is 0.67, not 1.0. The model learns to be uncertain where humans are uncertain, which is exactly what a triage system needs.
  • Report a ceiling. State the inter-annotator agreement alongside model accuracy so nobody mistakes label noise for model error.

Extreme imbalance, and how to sample

At a 0.2% base rate, a randomly sampled training batch of 1,024 posts contains about two positives. For a category at 0.001%, a batch of a million contains about ten.

The standard approach: keep every positive you have, and downsample negatives to a workable ratio — often somewhere between 1:10 and 1:100. Training becomes tractable and the positives are no longer drowned.

The consequence, which must be stated: the model's output probabilities are now on a shifted scale. A model trained at 1:10 when reality is 1:500 will output probabilities far too high. For a triage system that ranks by score, that is tolerable. For anything consuming the probability as a number, apply the correction from the ad click data lesson.

Two refinements that matter more than the ratio:

  • Stratify negatives. Random negatives are overwhelmingly trivial — holiday photos. Include a deliberate share of near-miss negatives: posts that were reviewed and cleared, posts that were reported and found fine, posts semantically close to positives. These are the hard negatives of Step 3: data and labels, and they are what teaches the model the boundary rather than the category.
  • Preserve the time distribution. Harm types have sharp temporal structure — a scam campaign runs for three weeks and vanishes. Sampling uniformly across two years produces a training set describing threats that no longer exist.

Active learning

With a fixed review budget, choosing which posts to label is as important as the model.

  • Uncertainty sampling. Send posts scoring near the decision boundary. They carry the most information per label.
  • Disagreement sampling. Where two models — or the model and a rule — disagree.
  • Coverage sampling. Deliberately label in regions of the feature space with few labels: new languages, new formats, new communities.
  • Keep the random slice. Uncertainty sampling alone leaves permanent blind spots, because a model is confidently wrong exactly where it has never seen anything. The random sample is the only thing that finds those.

A reasonable split of a fixed budget: 60% uncertainty, 20% coverage, 20% random.