Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Street View blurring: framing, the error asymmetry and labels
Prompt: "Design a system that automatically blurs faces and licence plates in street-level imagery before it is published."
Everything about this design follows from two facts established in the first two minutes: it is a batch problem, and the two error types have wildly different costs. This lesson establishes both, turns the second into a metric you can ship on, and then deals with the largest single cost in the project — the labels.
The clarifying questions
- Which object classes? Human faces and vehicle licence plates, at minimum. Do we also need house numbers, or people's bodies where they are identifiable without a face? Each class added is a separate annotation budget.
- What is the image volume and format? Single photos, or high-resolution panoramas that have to be tiled?
- Batch or real-time? This is the load-bearing question. Imagery is captured by vehicles, uploaded, processed, and published days later. There is no user waiting. That single answer removes the latency budget from the design and replaces it with a throughput and cost budget, which changes almost every downstream choice.
- What miss rate is acceptable? Ask directly. The answer will be "as close to zero as possible", and the follow-up — "what would you accept for the last percent?" — is the real conversation.
- Is there a takedown path? Users must be able to report a missed face. That path exists regardless, and its existence changes what the model has to achieve alone.
The constraints we design against
Invented throughout. The product is StreetLens, from an invented mapping company called Cartogram.
| Constraint | Value |
|---|---|
| Capture volume | ~100 million panoramas per month |
| Panorama size | 8000 × 4000 pixels, tiled into 12 overlapping crops for inference |
| Inference calls | ~1.2 billion crops per month |
| Turnaround | Published within 7 days of capture |
| Miss target | Fewer than 1 unblurred face per 10,000 published panoramas |
| Regions | 40 countries, with different legal regimes |
Framing it as machine learning
- Input: one image crop, roughly 1300 × 1300 pixels.
- Output: a list of bounding boxes, each with a class (face or plate) and a confidence score between 0 and 1.
- Objective: find every face and plate. Extra boxes are acceptable; missed ones are not.
The task type is object detection. Not classification — "does this image contain a face" does not tell you where to blur. Not segmentation — a pixel-exact mask is more than blurring needs, and costs far more to annotate and to run. Bounding boxes are exactly the right output shape, which is a satisfying and unusual position to be in.
The baseline first
Before any neural network: a classical face detector — the cascade-of-simple-features approach that has shipped in cameras for two decades — plus a plate detector built from edge density and aspect-ratio heuristics on high-contrast rectangular regions.
This runs on a CPU, costs almost nothing, and works. On frontal, well-lit, reasonably large faces it will find most of them. It fails on profiles, on partial occlusion, on faces smaller than about 30 pixels, and in low light — which is a large fraction of street imagery.
Building it anyway is worth two days, because it produces the annotation bootstrap (see Data and labels), the evaluation harness, and the pipeline, and it gives a defensible number for what the neural model has to beat.
Blur is not the only decision
One design note that separates a thought-through answer from a recited one: because this is a batch pipeline with no user waiting, you have a third option besides "blur" and "do not blur". You can route to a human. That option changes the whole metric discussion in the metrics lesson, and it exists only because there is no latency constraint.
Metrics, and the asymmetry that defines the system
Most classification problems ask you to balance precision and recall. This one does not. The two errors are not comparable, and the whole design falls out of saying so.
The cost of each error
A false negative is a published photograph of an identifiable person who did not consent. Depending on jurisdiction it is a privacy violation with regulatory exposure, it generates a takedown request that costs staff time, and it damages trust in the product. In some jurisdictions there is statutory liability. Treat this as expensive and partly unbounded.
A false positive is a blurred patch of wall, a smudged shop sign, or an over-blurred car window. The image is slightly worse. Nobody is harmed. A user might notice.
There is no exchange rate that makes these equal, and F1 — which assumes exactly that they are equal — is therefore the wrong metric here. Saying that plainly is one of the highest-value sentences available in this round.
Setting the threshold from a cost ratio
The model outputs a confidence per box. You choose a threshold above which you blur. That choice is where the asymmetry gets encoded.
Work it as an explicit cost calculation. Invented figures, used to show the method:
- Cost of a missed face: £400 in expected takedown handling, legal review, and reputational cost.
- Cost of an unnecessary blur: £0.02 in image-quality degradation, amortised.
- Ratio: 20,000 to 1.
The optimal threshold sits where the marginal expected cost of the two errors is equal. With a ratio that lopsided, the threshold goes very low — you blur anything with more than a few percent confidence. In practice you stop lowering it not because precision becomes bad but because the blurred area becomes visually unacceptable, and that is the real constraint.
So the metric is not F1 and not accuracy. It is:
Maximise recall subject to a cap on total blurred image area.
State the metric in that form. It is precise, it encodes the product constraint, and it is directly measurable.
mAP and IoU, explained
Detection metrics need a rule for when a predicted box counts as correct.
mean Average Precision (mAP) is the standard summary. For each class, sweep the confidence threshold from high to low, plot precision against recall, take the area under that curve (average precision), and average across classes. Usually reported at a stated IoU threshold — mAP@0.5 — or averaged over several.
For this system, mAP is useful for comparing models during development and is not the shipping metric, for two reasons. It averages over thresholds, and you have already decided to operate at a very low one. And it weights precision, which you have decided you care much less about.
The metrics that ship
| Metric | Definition | Target |
|---|---|---|
| Recall@IoU 0.5, faces | Fraction of ground-truth faces detected at the operating threshold | ≥ 99.5% |
| Recall on small faces (< 32 px) | Same, restricted to the hard slice | Reported separately — the aggregate hides it |
| Recall by region | Same, per country | Within 1 point of the global figure |
| Blurred area fraction | Blurred pixels ÷ total pixels | ≤ 1.2% |
| Human review volume | Boxes routed to a person per 1,000 panoramas | ≤ 40, on cost grounds |
| Takedown rate | Reported misses per million published panoramas | The true north star, arriving weeks late |
The takedown rate is the honest measure and it is unusable for iteration, because it arrives weeks after publication and only counts misses somebody noticed. Recall on the held-out annotated set is the proxy you actually optimise. Keeping both on the same dashboard, and watching whether they move together, is the discipline.
Data and labels
Every label here is a rectangle drawn by a person. That makes annotation the largest single cost in the project and the place where design effort pays back most.
The annotation cost
Drawing a tight box around every face and plate in a cluttered street scene takes a trained annotator roughly 20–40 seconds per image, and busy urban scenes take longer. At 30 seconds and an average of six objects per image, one million annotated images is on the order of 8,300 annotator-hours. That is a budget line requiring approval, not a task you assign on a Friday.
Three ways to reduce it, in order of value:
- Bootstrap with the classical baseline. Run the baseline detector from the framing lesson over a large sample, have annotators correct its output rather than draw from scratch. Correcting is roughly three times faster than drawing. The risk: annotators anchor on the model's output and miss what it missed, so a fraction of images must still be annotated blind, and that fraction is how you measure the anchoring bias.
- Active learning. Annotate where the model is uncertain rather than at random. After the first model exists, score a large unlabelled pool and send the low-confidence and high-disagreement images for labelling. Uniform random sampling wastes most of the budget on easy images the model already handles.
- Synthetic and composited data. Paste face crops into street scenes at controlled scales, angles, and lighting. Cheap and unlimited, and it never fully matches real capture conditions — use it to supplement the rare slices, not as the bulk.
The imbalance, which is spatial
This problem's class imbalance has an unusual shape. In a typical street panorama, faces and plates together occupy well under 0.5% of the pixels. A detector evaluates many thousands of candidate regions per image and nearly all of them are background.
If every candidate region contributes equally to the loss, the background term swamps everything. The model converges to "predict background everywhere", which achieves a superb loss and detects nothing. The training lesson covers the two standard fixes.
Augmentation for the conditions that actually occur
Augmentation here is not generic regularisation. Each one targets a real capture condition, and the list should be derived from looking at failures.
| Augmentation | The real-world case it covers |
|---|---|
| Scale jitter, 0.3× to 2× | Faces range from 12 to 400 pixels depending on distance |
| Rotation ±20° | Camera tilt on uneven road surfaces |
| Brightness and gamma shifts | Dawn, dusk, harsh midday sun, tunnels |
| Motion blur | Capture vehicle at 40 km/h |
| Random occlusion patches | Faces behind poles, railings, other people |
| Compression artefacts | Imagery stored lossily before processing |
| Weather overlays (rain, haze) | Regions and seasons with poor visibility |
Horizontal flip is safe for faces and not safe for licence plates, whose characters have orientation. That is the kind of detail worth mentioning: it shows the augmentation list was reasoned about rather than copied.
Sourcing the rare cases
Aggregate recall is easy. The last half-percent lives entirely in slices, and each needs deliberate sourcing.
- Small and distant faces. The largest single category of miss. Over-sample high-density pedestrian areas.
- Partial and occluded faces. Behind windscreens, in crowds, half out of frame.
- Faces in reflections. Shop windows and car bodies. Genuinely identifiable and genuinely hard.
- Non-standard plates. Formats differ by country: dimensions, colours, character sets, motorcycle plates, temporary plates.
- Faces on posters and advertisements. Arguably not a privacy issue and frequently blurred anyway. Decide the policy explicitly, then annotate consistently with it — inconsistent annotation on this one category can cost real recall elsewhere.
- Demographic and regional coverage. Detection performance varies with skin tone, facial hair, headwear, and lighting interactions. If the annotated set under-represents a group, the model's recall on that group will be lower and the aggregate metric will not show it. Sample the annotation pool to cover regions and appearances deliberately, and evaluate per slice. This is a requirement of the system, not an ethical footnote — a system that reliably blurs some people and not others fails at its actual job.