Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Street View blurring: the batch pipeline, human review and monitoring
No user is waiting. The design problem is moving 1.2 billion inference calls a month through accelerators at acceptable cost, and deciding what happens to the boxes the model is unsure about.
Then the system goes live, and a second gap opens. The model was measured on annotated data. Production is measured by what people report, and those two numbers can drift apart for months before anyone notices. This lesson covers the pipeline, the human queue that makes it safe, and the monitoring that keeps it honest.
The pipeline
Six stages, each independently scalable:
- Ingest. Capture vehicles upload panoramas to object storage. A manifest row per panorama enters a work queue.
- Tile. Each 8000 × 4000 panorama is cut into 12 overlapping crops. Overlap is 15% so a face on a tile boundary appears whole in at least one tile — a detail that costs 15% more compute and prevents a whole class of edge misses.
- Detect. Workers pull batches of crops, run the detector on an accelerator, write boxes back.
- Merge and suppress. Boxes from all 12 tiles are mapped back to panorama coordinates and non-maximum suppression is applied across tile boundaries to remove duplicates from the overlap regions.
- Route. High-confidence boxes go straight to blurring. Boxes in the uncertainty band go to the human queue. The panorama is held until every box is resolved.
- Blur and publish. Apply a Gaussian blur, generously padded beyond the box — a few pixels of margin costs nothing and covers localisation error. Write the published image; retain the original in restricted storage for reprocessing.
Throughput and cost
| Quantity | Value |
|---|---|
| Panoramas per month | 100,000,000 |
| Crops per panorama | 12 |
| Inference calls per month | 1,200,000,000 |
| Detector throughput per accelerator | ~25 crops/second |
| Accelerator-hours per month | ~13,300 |
| Sustained accelerators (24/7) | ~18 |
| Provisioned for burst (capture is seasonal) | ~60 |
Batching and utilisation
Accelerators are efficient only when fed. Three practical points:
- Batch crops, not panoramas. Fixed-size crops batch cleanly; variable-size panoramas do not.
- Decode on CPU, ahead of time. Image decoding is CPU work and will starve the accelerator if done inline. Use a prefetch pipeline with a CPU-to-accelerator ratio tuned so utilisation stays above 80%.
- Use reduced precision. Half-precision or 8-bit quantised inference typically gives a substantial throughput gain for a small accuracy cost. Measure the recall change on the small-face slice specifically, because that is where quantisation hurts first.
The human review queue
The queue exists because the model has an uncertainty band, and because with no latency constraint a human is a legitimate part of the inference path.
What enters: boxes scoring between 0.05 and 0.35, plus whole panoramas flagged by a heuristic — for example, a busy pedestrian scene where the detector found unusually few faces relative to the number of detected bodies.
Prioritisation. Not first-in-first-out. Order by expected harm: panoramas in dense residential areas ahead of empty motorways, and images from regions with stricter privacy law first. Reviewers see one panorama with candidate boxes overlaid and answer a single question per box: face, plate, or neither.
Throughput. At roughly 8 seconds per box, one reviewer handles about 450 boxes an hour. Forty boxes per 1,000 panoramas over 100 million panoramas is 4 million reviews a month — around 8,900 reviewer-hours, which is a team of roughly 55 people. That number is the reason the uncertainty band is a business decision.
When in doubt, blur. If the queue backs up beyond the 7-day publication deadline, the overflow policy is to blur everything in the band rather than publish unreviewed. Degrade toward the cheap error, never toward the expensive one.
The monitoring signals
Once imagery is publishing, watch for the gap between annotated-set performance and what people report.
| Signal | Cadence | What a change means |
|---|---|---|
| Boxes per panorama, by region | Daily | A sudden drop means a capture, tiling, or model regression |
| Blurred area fraction | Daily | Rising means over-detection; falling means the opposite |
| Human-queue volume | Daily | Rising means the model is less certain — often the first sign of drift |
| Reviewer agreement rate with the model | Weekly | Falling means the model's confidence is miscalibrated |
| Recall on a fixed golden set | Per model release | The controlled comparison |
| Takedown requests per million published | Weekly | The ground truth, arriving weeks late |
The human queue volume is the best early indicator available. It needs no labels, it moves within a day, and it responds to input drift — new camera hardware, a new country, a seasonal change in lighting — before the takedown rate can possibly react.
Takedown requests
A person reports an unblurred image of themselves. The system must:
- Blur the reported region within a stated turnaround, without waiting for a model update.
- Record the case as a labelled false negative and add it to the training set. Reported misses are the single highest-value training examples in the whole system, because they are real production failures on real data.
- Search for near-duplicates — the same street, the same capture run — because if one frame missed a face, the adjacent frames very likely did too. Reprocess that whole run.
Point 3 is the one candidates miss. A single report is evidence about a batch, not about one image.
Geographic and demographic performance differences
This belongs in the design, not in a closing caveat.
Face detection performance varies with skin tone, lighting interaction, facial hair, headwear, and image quality — and models trained on unevenly sampled data reliably inherit those gaps. Published audits of commercial face-analysis systems have found substantial accuracy differences across demographic groups; the specific figures vary by study and by system, and any number quoted here would need checking against a current source. The direction of the finding is well established enough to design against.
For a blurring system the consequence is direct and severe: lower recall for a group means that group's faces are published unblurred more often. The system fails at protecting the people it protects least well, which is the opposite of what a privacy system should do.
Concrete requirements, all of which belong in the design:
- Stratified evaluation sets covering skin tone, age, headwear, and region, each large enough for a meaningful recall estimate.
- Per-slice recall reported at every model release, with a release gate: no slice more than 1 point below the global figure.
- Annotation sampling that deliberately covers under-represented slices, rather than sampling uniformly from capture volume, which mirrors wherever the vehicles happen to drive.
- A rule that aggregate recall never ships alone. A model that raises global recall by 0.3 points while dropping one slice by 2 points is a regression.
Extensions
New object classes. House numbers, campaign posters, identifiable clothing or uniforms. Adding a class means new annotation and retraining, but the pipeline is unchanged — a good sign that the architecture was right.
Video rather than stills. Two changes. Frames are correlated, so tracking a detection across frames recovers misses: a face detected in frames 10–14 and missed in 12 can be interpolated. That is a large recall gain for modest cost. And temporal consistency matters visually — a blur that flickers on and off is worse than a steady one, so smooth the blur region across frames.
Continuous reprocessing. Old imagery was processed by old models. Periodically re-running the current detector over the archive finds misses nobody reported. Whether the recall gain justifies reprocessing billions of images is a cost decision, and the honest answer is to reprocess the highest-density regions and leave motorways alone.