Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Ad click prediction: serving inside the auction and calibration monitoring
Twenty milliseconds to score 800 candidate ads, inside a page load that is already happening. This is the tightest budget in the course, and it determines everything about the served model.
Once serving, the failure signal is unusual. In most systems, degraded quality is the symptom. Here, degraded calibration is the symptom, and it appears before anything else does. This lesson covers the auction path, continuous training, and the monitoring that catches mispricing early.
The full auction path
| Stage | Budget | Note |
|---|---|---|
| Request parse, user identification | 2 ms | |
| Targeting retrieval: eligible ads | 6 ms | Inverted index over targeting predicates |
| Budget and pacing filter | 2 ms | Remove ads out of budget or ahead of pace |
| Feature fetch (user + context) | 5 ms | One batched call; heavily cached |
| Click model scoring (800 ads) | 20 ms | 25 microseconds per ad |
| Auction: rank by bid × pCTR, apply reserve | 3 ms | |
| Ad quality and policy checks | 4 ms | |
| Response and logging | 3 ms | |
| Network and margin | 5 ms | |
| Total | 50 ms |
Twenty-five microseconds per ad is the constraint that determines everything about the model.
How to fit in the budget
Batch the scoring. All 800 candidates share the same user and context, so user features are fetched once and the user-side computation is done once. Only the ad-side and cross-feature parts vary. If the model is structured so the user representation is computed once and combined cheaply with each ad's representation, the marginal cost per ad collapses. That is the two-tower idea (Step 4: choosing a model) applied for latency rather than for retrieval — with the caveat that whatever cross features you want must still be computed per pair.
Compress the model. Four techniques, in order of usual value:
- Quantisation. Store weights and embeddings at 8-bit rather than 32-bit. Roughly 4× less memory and meaningfully faster; accuracy cost is usually small but must be measured on the calibration plot, not only on ROC-AUC.
- Distillation. Train a small model to reproduce the large model's outputs. Serve the small one. Often retains most of the quality at a fraction of the cost, and the student can be trained to match probabilities rather than labels, which preserves calibration.
- Pruning. Remove near-zero weights and unused embedding rows. Long-tail identifiers with almost no traffic can be dropped entirely into a shared "rare" bucket.
- Reduced embedding dimensions for low-importance feature fields.
Cache aggressively. User features change slowly within a session, so fetch once and reuse for every request in that session. Ad features change slowly and are shared across all users, so keep them in process memory. Some (user, ad) scores can even be cached briefly where context has not changed, though this must not extend across meaningfully different page contexts.
Fail open. If the model times out, fall back to the historical click-through rate baseline from the framing lesson. An auction that runs with worse predictions is better than a page with no ads — and the fallback being calibrated is exactly why that baseline stays in the system.
Continuous training
Ad performance moves fast. A creative launched this morning, a campaign that started at noon, a news event shifting what people engage with — a model trained on last week does not know about any of it.
Three cadences, and the trade-offs:
| Cadence | Freshness | Operational cost | Risk |
|---|---|---|---|
| Daily batch retrain | Up to 24 h stale | Low | Misses fast-moving creatives entirely |
| Hourly incremental | Up to 1 h stale | Medium | Needs a reliable streaming feature pipeline |
| Continuous online learning | Minutes | High | A bad batch of data damages the live model immediately |
The usual production shape: a stable base model retrained daily on a full window, plus a fast-adapting layer — often the per-creative and per-campaign counting features, or a small correction model — updated every few minutes from the stream. The base model carries the general structure; the fast layer tracks what changed this morning.
Continuous learning needs guard rails, and naming them matters: validate every update against a held-out recent window before it takes effect, bound how far parameters can move in one update, and keep the ability to roll back to a checkpoint from any point in the last day. An online learner with no validation gate will eventually be poisoned by a broken feature or an anomalous traffic hour.
Calibration drift as the leading indicator
The predicted-to-actual click ratio from the metrics lesson is the single most valuable production metric in this system. Sum predicted probabilities over an hour, divide by observed clicks. It requires no labelled evaluation set, it updates hourly, and it is sensitive to almost every failure mode this system has.
What moves it, and what each cause looks like:
| Cause | Signature |
|---|---|
| A feature pipeline broke, values now default | Sharp step change, often within one hour, usually in one segment |
| Traffic mix shifted (a country, a placement, a device grew) | Gradual drift; global ratio moves while per-segment ratios are stable |
| Advertiser mix shifted (a large campaign started) | Drift concentrated in one category |
| Model retrained with a different downsampling rate | Sharp step at deploy time, ratio off by a clean factor |
| Genuine concept drift | Slow, persistent drift across all segments |
The diagnostic is always the same: look at the ratio per segment. If it moved globally but is stable in every segment, the mix changed and the model is fine. If it moved in one segment, that segment has a problem.
Both axes are percentages: predicted click probability against observed click rate in that bucket. Read three things. The healthy model tracks the diagonal within noise across the whole range. The drifted model sits below it everywhere — it over-predicts, so observed clicks fall short of predictions. And the gap widens with the predicted value, which means the mispricing is worst exactly on the high-value impressions where the money is. Values are illustrative.
An ordering metric cannot see any of this: the drifted curve is monotonic, so its ROC-AUC is close to the healthy model's.
Position and selection bias
Position bias. An ad in the first slot is clicked far more than the same ad in the fourth. Train on raw clicks and the model partly learns position. The fix is the same as in the video search data and features lesson: include position as a training feature and set it to a fixed reference value at serving time, so every candidate is scored as if it were in the same slot.
Selection bias. You observe clicks only on ads that won auctions. An ad that always loses generates no data, so it never gets a chance to prove itself. Log the win probability so offline evaluation can be reweighted, and reserve an exploration budget for under-served ads.
Ad fatigue
The same creative shown repeatedly to the same user loses effectiveness. The click rate for a user's fifth exposure to a creative is materially lower than for the first — a well-documented industry effect, though the magnitude varies enormously by format, placement, and category, so any specific figure needs measuring rather than quoting.
Handle it as a feature, not a rule: times_this_user_saw_this_creative and hours_since_last_exposure let the model learn the decay curve, which differs by category, and the resulting predictions handle frequency naturally inside the auction. Add a hard frequency cap on top as a user-experience protection, because there is a point past which repetition is an annoyance regardless of what the model predicts.
Budget pacing
An advertiser with a £5,000 daily budget should not spend it by 09:00. Pacing decides which auctions to enter over the day.
It interacts with prediction in a way worth naming: pacing is usually implemented as a multiplier on the bid or as a probabilistic participation gate, and both change which impressions the advertiser wins. That changes the data the model sees for that advertiser, which changes future predictions. Pacing and prediction form a control loop, and an unstable pacing controller shows up as oscillating predicted click rates for affected campaigns.
Cold start for new ads
A new creative has no history, so its counting features are empty. Three mitigations:
- Hierarchical fallbacks. Fall back through creative → campaign → advertiser → category → global, at each level using the most specific estimate with enough data, smoothed toward the next level up.
- Content features. The creative's image and text can be embedded, so a new creative is scored on what it looks like and says rather than on nothing.
- Exploration allocation. Guarantee new creatives a slice of impressions so they generate data. Without it, a new creative with a conservative initial estimate loses every auction and never gets the impressions that would correct the estimate — the same self-fulfilling failure as in Section 6.