Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Street View blurring: the detector and its training


A detector has to answer "is there an object here" and "exactly where" at the same time, at every position and every scale in the image. The architectures differ in how they split that work.

For StreetLens — 1.2 billion crops a month, no user waiting, and a missed face costing 20,000 times an extra blur — the choice is driven by cost rather than latency, and the asymmetry from the metrics discussion reaches all the way into post-processing parameters. Training then has its own trap: two losses running at once and a background class that will drown both if you let it.

The two families

Two-stage detectors. A first network proposes a few thousand regions that might contain something. A second network classifies each proposal and refines its box. More accurate, especially on small objects, because the second stage sees a cropped and resized version of each region at higher effective resolution. Slower, because the second stage runs many times per image.

Single-stage detectors. One network predicts class and box offsets directly at every position of a feature grid, in one pass. Faster — often three to five times — and historically weaker on small objects, though the gap has narrowed considerably with better training techniques.

Two-stageSingle-stage
Relative throughput1×3–5×
Small-object recallBetterGood, and improving
Implementation complexityHigherLower
Cost at 1.2 billion crops/monthHighSubstantially lower

The choice for StreetLens, and why cost decides it

There is no latency constraint. There is a very large throughput constraint, and in a batch pipeline throughput is cost.

Work it. At 1.2 billion crops per month:

  • A single-stage detector at 25 crops per second per accelerator needs about 13,300 accelerator-hours per month.
  • A two-stage detector at 7 crops per second per accelerator needs about 47,600.

At an illustrative £1.60 per accelerator-hour that is roughly £21,000 versus £76,000 a month. The two-stage detector must buy about £55,000 a month of value in additional recall to be worth it.

The recommendation: a single-stage detector with a feature pyramid — multi-scale feature maps so small and large faces are each detected at an appropriate resolution — run at higher input resolution than default, because small faces are where recall is lost and resolution is the cheapest way to buy it back. Then route low-confidence detections to human review, which recovers accuracy far more cheaply than doubling the model cost across every image.

That trade — cheap model plus targeted human review, rather than expensive model everywhere — is the central design decision of this case study.

Anchor boxes

Non-maximum suppression

A detector fires on several nearby anchors for the same face. Left alone it returns eight overlapping boxes for one person.

Non-maximum suppression (NMS) cleans this up:

  1. Sort all boxes for a class by confidence, highest first.
  2. Take the highest. Keep it.
  3. Discard every remaining box whose IoU with it exceeds a threshold (commonly 0.5).
  4. Repeat with the next surviving box.

The threshold is a genuine trade-off. Set it low (0.3) and boxes are aggressively merged — two people standing shoulder to shoulder may collapse into one detection, and you have created a false negative during post-processing. Set it high (0.7) and you keep duplicates, which for blurring is harmless because overlapping blurs are still blurs.

Given the asymmetry, use a high NMS threshold. Duplicate blurs cost nothing; a merged detection that misses the second person costs a privacy incident. This is the asymmetry from the metrics lesson reaching into a post-processing parameter, and it is exactly the kind of connection interviewers reward.

Detection produces overlapping boxes; NMS keeps one per objectraw detections0.910.870.790.880.72one object, several boxes, each with a confidencesorted, highest first0.91IoU 0.86IoU 0.790.88IoU 0.81keep the top box, then suppress anything overlapping itabove the IoU thresholdafter NMS0.910.88one box per objectThe IoU threshold is the tunable: too low and two genuinely adjacent faces get merged into one; too high and every face keeps three boxes.For blurring, recall matters far more than precision — a missed face is a privacy incident, an extra blur is a cosmetic complaint.
NMS is a post-process, not part of the model — which is why its threshold can be retuned without retraining anything.

Training: the two-part loss

With the architecture chosen, training has two losses running at once and a background class that will drown both if you let it.

Why background drowns the lossbgbgbgbgbgbgbgfacebgbgplatebgbgbgbgbgbgbgbgbgbgbgbgbgRoughly two positive anchors among tens of thousands.
Hard negative mining and focal loss both exist to stop easy background from supplying almost all of the gradient.

Every detector optimises a sum:

Text
total loss = classification loss + λ · localisation loss
  • Classification loss — for each anchor, is this background, a face, or a plate? Usually cross-entropy, and the part the imbalance attacks.
  • Localisation loss — for anchors matched to a real object, how far off are the four box coordinates? Usually a smooth L1 loss (quadratic near zero, linear in the tail) so a badly wrong box does not dominate the gradient. Losses computed directly on IoU are an increasingly common alternative and optimise the metric more directly.
  • λ balances the two. Localisation loss is computed only over matched anchors, so it sees far fewer terms; λ compensates.

Only matched anchors contribute localisation loss. Unmatched anchors contribute classification loss only — and there are tens of thousands of them.

The background problem, with numbers

One crop produces roughly 40,000 anchors. A busy street scene contains perhaps 12 faces and 6 plates. Even generously matching three anchors per object, that is 54 positive anchors against 39,946 negatives — a ratio of about 740 to 1.

Nearly all of those negatives are trivially easy: a patch of sky, a patch of tarmac. Each one individually produces a tiny loss, but 39,946 tiny losses sum to far more than 54 meaningful ones. The gradient points toward "say background more confidently", and training stalls with excellent loss and no detections.

Fix one: hard negative mining

Compute the loss for all negative anchors, sort them, and keep only the worst — typically enough to reach a 3:1 negative-to-positive ratio. Discard the rest entirely.

It works and it is a sorting step inside every training iteration, and it introduces a sensitive hyperparameter (the ratio) that interacts with batch size.

Fix two: focal loss

Focal loss is the recommendation here — no sorting step, no ratio to tune, and it uses every negative rather than discarding most of them. Keep γ around 2 and note that it needs a careful bias initialisation on the final layer, or early training diverges.

Evaluating on a held-out geography

Step 5: training said split by time. This problem needs a different cut: split by geography.

Take every image from three cities and hold them out entirely. Train on the rest. Evaluate only on the held-out cities.

The reason is specific. StreetLens captures a city by driving it, and consecutive frames along one street show the same parked cars, the same shopfronts, the same pedestrians from slightly different angles. Split at random and near-duplicate frames land on both sides. The model memorises specific vehicles and specific faces, and recall on the test set is inflated by several points — an invented but very plausible figure would be a random split reporting 99.3% recall where a city-held-out split reports 96.1%.

Which number is real? The one measuring performance on a city the model has never seen, which is the situation in production every single day.

Add a second held-out slice on a different continent, to measure generalisation across plate formats, vehicle types, streetscapes, and demographics. If recall drops sharply there, you have found a coverage gap in the training data, and that is a fairness finding as much as an accuracy one.