Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Event recommendations: framing and the registration-attendance gap


Prompt: "Design a system that recommends local events to users."

This case study was chosen because its constraints break the template Section 6 (Video Recommendation System) established. Every item is brand new, expires on a fixed date, and is useless to anyone more than an hour's travel away. The techniques that carried video recommendation barely function here.

This lesson frames the problem and sets its metrics. The metrics have an unusual property: the outcome you care about happens in the physical world, days after the prediction, and a large fraction of the people who said yes do not turn up.

Why the feed template does not transferWhat breaks here• Every event is new and then expires• No session history to embed• Attendance is days after the clickWhat carries it instead• Distance, time, price as direct features• Category affinity from past attendance• Trees over hand-built features
The catalogue turns over completely, so the system cannot learn item embeddings and must learn from item attributes instead.

The clarifying questions

  1. How far ahead do we recommend? Tonight, this weekend, or the next three months? Each is a different candidate pool and a different notion of relevance.
  2. What radius? A 5 km walk, a 40 km drive, or anywhere in the country for a major event? Distance is a hard filter, not a soft feature.
  3. What is the goal — registration, attendance, or ticket purchase? They diverge more than you would expect, and the metrics lesson makes that gap the metric story.
  4. How large is the catalogue, and what is the turnover?
  5. Are group and social signals available? Whether a user's connections are attending is one of the strongest features in this problem, if the platform has that data.

The constraints we design against

Invented. The product is Assembly, a platform for local events — meetups, gigs, classes, markets.

ConstraintValue
Users8 million monthly active
Live events at any moment~400,000 globally, upcoming within 90 days
Events per user after location and date filters~600
New events per day~35,000
Events expiring per day~35,000
Median event lifetime from listing to occurrence18 days
Requests~2,000 per second at peak
Latency budget300 ms p99

The third row is the number that changes the design. After filtering by location and date, the candidate pool per user is about 600 items — not 50 million. This is the smallest candidate space in the course and it makes two-stage retrieval unnecessary. The serving lesson develops the consequence.

Framing it as machine learning

  • Input: a user with profile and history, a context (current location, time, day of week), and the set of events within the filter window.
  • Output: a ranked list of 20 events.
  • Objective: maximise the probability the user registers and attends.

Task type: ranking, over an already-small candidate set. A pointwise model predicting P(register) per event, sorted, is entirely adequate — there is no retrieval stage to build.

The three constraints that break the Section 6 template

1. Severe cold start, permanently. In video recommendation, a new item is cold for a few hours and then has engagement history. Here, the median event exists for 18 days and is attended once. It never accumulates a history. Every item is cold at the moment it matters most, and there is no steady state where collaborative signals become available.

Collaborative filtering asks "which users who liked this also liked that". With items that appear, are attended once, and vanish, the co-occurrence matrix is almost empty. The technique that carried Section 6 does not apply.

2. Time decay, in two directions. An event has a fixed date. Its relevance is not monotonic: a concert three months away is too distant to act on, one this Friday is ideal, one that started an hour ago is worthless. Value peaks somewhere in the middle and drops discontinuously to zero at the start time. Recommendation systems usually deal with monotonic freshness decay; this is a curve with a peak and a cliff.

3. A hard location filter. A brilliant event 400 km away is not a slightly worse recommendation — it is a wrong one. Distance is not a feature to be traded against relevance; it is a precondition. Getting this right before the model, rather than hoping the model learns it, is the serving decision in the serving lesson.

What carries the system instead

Because behavioural history cannot, three other things do:

  • Content features — category, description, format, price, time of day.
  • Organiser history — the organiser is not new even though the event is. An organiser who has run 40 well-attended photography walks is strong evidence about their 41st. This is the single most valuable substitution available.
  • Category priors and user-stated preferences — declared interests, past categories attended, and the demographics of who typically attends a category.

This inverts the usual recommendation design. In Section 6, content features were a fallback for the cold-start minority. Here, they are the system, and behavioural signals are the occasionally-available bonus.

The baseline first

Rank by distance ascending within the date window, breaking ties by attendee count. No model, no training data, one day of work.

For a user in a dense city this is a real product. Nearby and popular is a strong ordering, and it will beat a badly-built model. Any personalised system has to beat "close and popular", and stating that plainly sets the bar honestly.

Metrics

This problem has an unusual property: the outcome you care about happens in the physical world, days after the prediction, and a large fraction of the people who said yes do not turn up.

The label arrives after the event does0123456701234567recommendedregisteredevent dayDays. The true label is attendance, known only after day 6.
Registration is a fast proxy for a slow outcome, and a large share of the people who register never turn up.

Offline metrics

Offline evaluation replays past sessions: given the events a user was shown on a given day, would the new model have ranked the one they registered for higher?

MetricNote
nDCG@20 on registrationsThe headline. Graded by outcome: attended > registered-not-attended > viewed > ignored
precision@5Most users look at the top few only; the top of the list is where the value is
nDCG@20 on attendanceThe honest version, and it needs check-in data
Calibration of P(register)Matters if the score feeds notification decisions — see Monitoring and follow-ups

The graded relevance scale is worth building deliberately. Treating "registered" as the only positive throws away the distinction between a registration that became an attendance and one that did not, and that distinction is the most interesting signal in the problem.

Online metrics

MetricLayerReading
Event page view rate from recommendationsImmediateAre the cards attractive?
Registration rateSessionThe usual headline
Attendance rateDelayed, daysDid they actually go?
The registration-to-attendance gapDelayedThe metric this problem is really about
Repeat registration within 30 daysRetentionDid the experience justify the recommendation?
Guard: notification opt-out rateImmediateOver-recommending is punished here more than elsewhere

The gap between registering and attending

Free events on invented Assembly figures see roughly 55–65% of registrants actually attend; paid events see 85–90%. Those specific numbers are invented, but the direction — that free registration is a much weaker commitment than paid — is well documented across the events industry and worth checking against a current source before publication.

The gap is the most informative signal in the system, and no click-based metric captures it.

Consider two events, both with a 12% registration rate from recommendations:

  • Event A: 85% of registrants attend. Effective yield 10.2%.
  • Event B: 40% of registrants attend. Effective yield 4.8%.

A system optimising registrations rates them identically. A system optimising attendance prefers A by more than two to one. The difference is real: Event B is generating registrations from people who liked the idea and were never going to travel across the city on a Tuesday night.

What drives the gap, and therefore what belongs in the model:

  • Distance. The strongest predictor. Registration is free; travelling 30 km is not.
  • Time of day and day of week. A Wednesday 19:00 event competes with the rest of a weekday evening.
  • Lead time. Registering three months ahead is a weaker commitment than registering two days ahead — the intention decays.
  • Price. Payment is commitment.
  • Weather, for outdoor events. Genuinely predictive and genuinely usable, since forecasts exist at the relevant horizon.

The design consequence: train on attendance where check-in data exists, and on registration elsewhere, weighting the registration-only examples lower. Or predict both as separate heads and combine, which is the Section 10 pattern applied at small scale.

The delayed-label problem

The label for a prediction made today arrives when the event happens — a median of 18 days later, sometimes 90. That has three consequences:

  1. Feedback is slow. A model shipped today cannot be evaluated on attendance for weeks. Registration is the fast proxy, and its bias toward low-commitment registrations is the price.
  2. Training data has a horizon. The most recent three months of data have incomplete attendance labels. Either wait, or train on registration for the recent window and attendance for the older one.
  3. Experiments must run long. An A/B test measuring attendance needs to run for the registration period plus the full lead time. Six weeks, not one.