Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Event recommendations: filter-first serving, seasonality and sparse regions


The distinctive serving idea in this case study is that the cheap deterministic filters run first, and they do more work than any model. That idea carries a lesson that transfers well beyond events: the two-stage architecture is a response to a constraint, not a law.

After launch, the system's failures are geographic and seasonal, which means aggregate dashboards hide almost all of them. The second half of this lesson covers how to see them and what to do — often with a product response rather than a better ranker.

Filter before you rank

The naive ordering is to score everything and let the model handle location. That is wrong for three reasons, and articulating them well is most of the value of this step.

It is wasteful. Scoring 400,000 events per request at 0.05 ms each is 20 seconds. Filtering to 600 first makes the model cost 30 milliseconds. The filters are exact predicate evaluations over indexed fields and cost a couple of milliseconds.

It is less correct. A model asked to learn "never recommend anything more than 50 km away" will learn it approximately. It will occasionally rank a distant event highly because every other signal is strong, and that is a visibly broken recommendation. A hard filter is exactly right, every time, and it is inspectable.

It is unmaintainable. Radius by user preference, date window by product surface, and sold-out exclusion all change on product timelines, sometimes weekly. Changing a predicate is a configuration deployment; changing what a model has learnt is a retraining cycle.

The general principle, and it transfers to every problem in this course: express hard constraints as filters, and soft preferences as model features. Distance beyond the radius is a constraint. Distance within the radius is a preference.

The reduction, stage by stage

StageRemainingCostMechanism
All live events globally400,000——
Within the user's radius~2,4002 msGeospatial index lookup (see geospatial indexing in System Design Interview)
Within the date window~6001 msRange predicate on start time
Not already registered, not hidden, not cancelled~5701 msSet difference against a user exclusion list
Eligibility: age-appropriate, capacity remaining, region-permitted~5402 msPredicate checks
Into the ranking model~540——
After ranking and diversity re-rank20——

From 400,000 to 540 in six milliseconds, with no model involved. The candidate-set reduction that Section 6 needed a trained retrieval model and an approximate nearest neighbour index to achieve, this problem gets from a geospatial index and two range predicates.

That is the teaching point of this case study: the two-stage architecture is a response to a constraint, not a law. When cheap exact filters reduce the candidate set enough, adding a learned retrieval stage adds cost, failure modes, and a recall ceiling for nothing. Recognising when the standard pattern is unnecessary is a stronger signal than applying it everywhere.

All events2,000,000everything in the catalogueGeographic filter8,000within 50 km — cheap, and removes99.6%Time filter1,200in the next 30 daysEligibility filter900not sold out, not already attended,age-appropriateRanking model900 scoredthe only expensive stageRe-rank +diversify20 showncategory spread, freshnessthe expensive stage runs on 900 items, not 2million — the filters bought that by a factor of 2,000
Filter order matters: the cheapest, most selective filter goes first, and geography is almost always both.

The rest of the path

Feature hydration. With only ~540 candidates, a single batched feature-store call is comfortable. Item and organiser features are cached aggressively since they are shared across users; user features are one lookup; cross features are computed in the request.

Ranking. The tree model over ~540 candidates. Boosting inference is fast and CPU-friendly — no accelerators needed, which is a cost advantage worth mentioning.

Re-ranking. Category diversity (at most 5 of 20 from one category), organiser cap (at most 2 from one organiser), a slot or two for events happening today, and an exploration slot for a new organiser with no history.

Latency budget, against 300 ms p99:

StageBudget
Request handling and user context10 ms
Filters (geospatial, date, exclusion, eligibility)6 ms
User feature fetch20 ms
Batched candidate feature hydration (~540)35 ms
Cross-feature computation15 ms
Tree model scoring30 ms
Re-rank and diversity8 ms
Response assembly12 ms
Network and margin45 ms
Total181 ms

Comfortable headroom, which is what you would expect from a problem with a 540-item candidate set. Spend the surplus on richer cross features rather than on a bigger model.

Monitoring

This system's failures are geographic and seasonal, which means aggregate dashboards hide almost all of them.

Recall at 10, sliced by density and season0.420.390.310.240.210.170.090.080.05SummerAutumnWinterDense citySuburbRuralThe aggregate number is dominated by dense cities.
A healthy global metric can hide a system that recommends nothing usable to a whole region for half the year.
SignalCadenceWhat it detects
Registration rate by cityDailyThe aggregate is dominated by three big cities; everywhere else fails invisibly
Attendance rate by cityWeeklySlow, and the honest measure
Events with zero impressionsDailyThe supply side. An organiser whose event was never shown will not list again
Candidate-set size distributionDailyUsers whose filters return under 20 events — the sparse-region problem
Feature null ratesHourlyWeather feed down, geospatial service down, social graph stale
Score distributionHourlyThe fastest label-free drift signal
Notification opt-out rateDailyOver-recommendation shows up here first

The last row deserves emphasis. Recommendations here are frequently delivered as notifications — "three events near you this weekend" — and a notification for a bad recommendation costs far more than a bad card in a feed. It is an interruption. Track opt-out and mute rates as first-class guard metrics, and set a confidence threshold for notification that is much higher than for in-feed display. The score's calibration matters for exactly this reason, which is why the metrics lesson listed it.

Seasonality

This problem has stronger seasonality than anything else in the course, in three layers:

  1. Weekly. Weekend events dominate. A model trained without day-of-week features will systematically misestimate weekday events.
  2. Annual. Outdoor events in summer, indoor in winter; festivals cluster; December is unlike November in every category.
  3. Irregular. Holidays, school terms, major sporting fixtures, and weather events shift behaviour sharply for days at a time.

Consequences for the design:

  • Train on at least 18 months so every season appears twice. Twelve months gives one observation per season, which is not enough to separate seasonality from trend.
  • Evaluate on the matching season. A model trained through October and tested on November will not tell you how it performs in June.
  • Include explicit seasonal features: month, week of year, days to the nearest public holiday, school-term flag, and a weather forecast.
  • Do not over-correct for a one-off. A single unusual week — a heatwave, a transport strike — should not reshape the model. Down-weight or exclude anomalous periods explicitly.

Sparse regions

A user in a small town may have eleven events within their radius in the next month. The ranking model is nearly irrelevant — the list is short enough to show whole.

The product response matters more than the model response:

  • Widen the radius automatically when the candidate count is below a floor, and label it in the interface ("nothing nearby this week — showing events within 60 km").
  • Extend the date window rather than the radius, where that fits the category.
  • Show online events, which have no distance constraint.
  • Shift the objective to supply. In a sparse region the highest-value action is not a better ranking; it is recruiting organisers. Surface "start an event" prompts and track region supply as a product metric.

Monitoring the candidate-set size distribution is what makes this visible. Aggregate registration rate will never reveal it, because sparse regions are a small share of traffic and a large share of the dissatisfaction.

Large events versus small relevant ones

A structural tension. A 4,000-person festival will out-perform a 12-person book club on almost every engagement metric, because more people are interested in it. Rank purely on predicted registration and the feed fills with large events, small organisers get no exposure, and the long tail of the catalogue dies — the popularity loop from the video recommendation monitoring lesson.

Three counterweights:

  1. Normalise by event capacity. Predict registration rate as a fraction of capacity rather than absolute count. A book club filling 12 of 12 places is performing better than a festival filling 4,000 of 20,000.
  2. Reserve slots. Guarantee some positions for small or new-organiser events, exactly as Section 6 reserves exploration slots.
  3. Optimise attendance value, not registration count. A small, well-matched event has a higher attendance rate and a higher chance of a repeat booking. Those are the outcomes worth having.

This is a marketplace health decision as much as a ranking one: without small events, organisers stop listing, and the catalogue that makes the product distinctive disappears.