Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Event recommendations: filter-first serving, seasonality and sparse regions
The distinctive serving idea in this case study is that the cheap deterministic filters run first, and they do more work than any model. That idea carries a lesson that transfers well beyond events: the two-stage architecture is a response to a constraint, not a law.
After launch, the system's failures are geographic and seasonal, which means aggregate dashboards hide almost all of them. The second half of this lesson covers how to see them and what to do — often with a product response rather than a better ranker.
Filter before you rank
The naive ordering is to score everything and let the model handle location. That is wrong for three reasons, and articulating them well is most of the value of this step.
It is wasteful. Scoring 400,000 events per request at 0.05 ms each is 20 seconds. Filtering to 600 first makes the model cost 30 milliseconds. The filters are exact predicate evaluations over indexed fields and cost a couple of milliseconds.
It is less correct. A model asked to learn "never recommend anything more than 50 km away" will learn it approximately. It will occasionally rank a distant event highly because every other signal is strong, and that is a visibly broken recommendation. A hard filter is exactly right, every time, and it is inspectable.
It is unmaintainable. Radius by user preference, date window by product surface, and sold-out exclusion all change on product timelines, sometimes weekly. Changing a predicate is a configuration deployment; changing what a model has learnt is a retraining cycle.
The general principle, and it transfers to every problem in this course: express hard constraints as filters, and soft preferences as model features. Distance beyond the radius is a constraint. Distance within the radius is a preference.
The reduction, stage by stage
| Stage | Remaining | Cost | Mechanism |
|---|---|---|---|
| All live events globally | 400,000 | — | — |
| Within the user's radius | ~2,400 | 2 ms | Geospatial index lookup (see geospatial indexing in System Design Interview) |
| Within the date window | ~600 | 1 ms | Range predicate on start time |
| Not already registered, not hidden, not cancelled | ~570 | 1 ms | Set difference against a user exclusion list |
| Eligibility: age-appropriate, capacity remaining, region-permitted | ~540 | 2 ms | Predicate checks |
| Into the ranking model | ~540 | — | — |
| After ranking and diversity re-rank | 20 | — | — |
From 400,000 to 540 in six milliseconds, with no model involved. The candidate-set reduction that Section 6 needed a trained retrieval model and an approximate nearest neighbour index to achieve, this problem gets from a geospatial index and two range predicates.
That is the teaching point of this case study: the two-stage architecture is a response to a constraint, not a law. When cheap exact filters reduce the candidate set enough, adding a learned retrieval stage adds cost, failure modes, and a recall ceiling for nothing. Recognising when the standard pattern is unnecessary is a stronger signal than applying it everywhere.
The rest of the path
Feature hydration. With only ~540 candidates, a single batched feature-store call is comfortable. Item and organiser features are cached aggressively since they are shared across users; user features are one lookup; cross features are computed in the request.
Ranking. The tree model over ~540 candidates. Boosting inference is fast and CPU-friendly — no accelerators needed, which is a cost advantage worth mentioning.
Re-ranking. Category diversity (at most 5 of 20 from one category), organiser cap (at most 2 from one organiser), a slot or two for events happening today, and an exploration slot for a new organiser with no history.
Latency budget, against 300 ms p99:
| Stage | Budget |
|---|---|
| Request handling and user context | 10 ms |
| Filters (geospatial, date, exclusion, eligibility) | 6 ms |
| User feature fetch | 20 ms |
| Batched candidate feature hydration (~540) | 35 ms |
| Cross-feature computation | 15 ms |
| Tree model scoring | 30 ms |
| Re-rank and diversity | 8 ms |
| Response assembly | 12 ms |
| Network and margin | 45 ms |
| Total | 181 ms |
Comfortable headroom, which is what you would expect from a problem with a 540-item candidate set. Spend the surplus on richer cross features rather than on a bigger model.
Monitoring
This system's failures are geographic and seasonal, which means aggregate dashboards hide almost all of them.
| Signal | Cadence | What it detects |
|---|---|---|
| Registration rate by city | Daily | The aggregate is dominated by three big cities; everywhere else fails invisibly |
| Attendance rate by city | Weekly | Slow, and the honest measure |
| Events with zero impressions | Daily | The supply side. An organiser whose event was never shown will not list again |
| Candidate-set size distribution | Daily | Users whose filters return under 20 events — the sparse-region problem |
| Feature null rates | Hourly | Weather feed down, geospatial service down, social graph stale |
| Score distribution | Hourly | The fastest label-free drift signal |
| Notification opt-out rate | Daily | Over-recommendation shows up here first |
The last row deserves emphasis. Recommendations here are frequently delivered as notifications — "three events near you this weekend" — and a notification for a bad recommendation costs far more than a bad card in a feed. It is an interruption. Track opt-out and mute rates as first-class guard metrics, and set a confidence threshold for notification that is much higher than for in-feed display. The score's calibration matters for exactly this reason, which is why the metrics lesson listed it.
Seasonality
This problem has stronger seasonality than anything else in the course, in three layers:
- Weekly. Weekend events dominate. A model trained without day-of-week features will systematically misestimate weekday events.
- Annual. Outdoor events in summer, indoor in winter; festivals cluster; December is unlike November in every category.
- Irregular. Holidays, school terms, major sporting fixtures, and weather events shift behaviour sharply for days at a time.
Consequences for the design:
- Train on at least 18 months so every season appears twice. Twelve months gives one observation per season, which is not enough to separate seasonality from trend.
- Evaluate on the matching season. A model trained through October and tested on November will not tell you how it performs in June.
- Include explicit seasonal features: month, week of year, days to the nearest public holiday, school-term flag, and a weather forecast.
- Do not over-correct for a one-off. A single unusual week — a heatwave, a transport strike — should not reshape the model. Down-weight or exclude anomalous periods explicitly.
Sparse regions
A user in a small town may have eleven events within their radius in the next month. The ranking model is nearly irrelevant — the list is short enough to show whole.
The product response matters more than the model response:
- Widen the radius automatically when the candidate count is below a floor, and label it in the interface ("nothing nearby this week — showing events within 60 km").
- Extend the date window rather than the radius, where that fits the category.
- Show online events, which have no distance constraint.
- Shift the objective to supply. In a sparse region the highest-value action is not a better ranking; it is recruiting organisers. Surface "start an event" prompts and track region supply as a product metric.
Monitoring the candidate-set size distribution is what makes this visible. Aggregate registration rate will never reveal it, because sparse regions are a small share of traffic and a large share of the dissatisfaction.
Large events versus small relevant ones
A structural tension. A 4,000-person festival will out-perform a 12-person book club on almost every engagement metric, because more people are interested in it. Rank purely on predicted registration and the feed fills with large events, small organisers get no exposure, and the long tail of the catalogue dies — the popularity loop from the video recommendation monitoring lesson.
Three counterweights:
- Normalise by event capacity. Predict registration rate as a fraction of capacity rather than absolute count. A book club filling 12 of 12 places is performing better than a festival filling 4,000 of 20,000.
- Reserve slots. Guarantee some positions for small or new-organiser events, exactly as Section 6 reserves exploration slots.
- Optimise attendance value, not registration count. A small, well-matched event has a higher attendance rate and a higher chance of a repeat booking. Those are the outcomes worth having.
This is a marketplace health decision as much as a ranking one: without small events, organisers stop listing, and the catalogue that makes the product distinctive disappears.