Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Event recommendations: cold start, features and gradient-boosted trees
Most recommendation systems have a cold-start problem. This one has a cold-start condition — it never ends, and the design has to be built for it rather than around it.
That condition drives everything in this lesson. With no interaction history to lean on, the whole system lives or dies on feature engineering, which is unfashionable and, for this problem, correct. And once the features are tabular and engineered, this is the case study where the honest answer is not a neural network — saying so is the point.
Why cold start is different here
| Video recommendation | Event recommendation | |
|---|---|---|
| Item lifetime | Years | 18 days median, then gone |
| Interactions per item | Thousands to millions | Tens to hundreds |
| Time to accumulate signal | Hours | Never enough |
| Can collaborative filtering work? | Yes, it is the core | Barely |
| Value of content features | Fallback for new items | The primary signal |
Run the arithmetic. 35,000 new events a day, 8 million monthly users, and a typical event attracting perhaps 60 registrations. The user-event interaction matrix has roughly 400,000 columns of which each holds a few dozen non-zero entries out of 8 million rows. That is a density around 0.0008%. Matrix factorisation on a matrix that sparse learns essentially nothing about individual items.
Worse, by the time an event has accumulated enough interactions to be modelled well, it is about to happen and there is little value left to extract.
What replaces collaborative filtering
Organiser history — the key substitution. The event is new; the organiser rarely is.
For an organiser with prior events, compute: number of events run, median attendance rate, median registration count, category consistency, rating from past attendees, and cancellation rate. An organiser with 40 well-attended photography walks tells you a great deal about their 41st, and none of it requires the new event to have any history.
This turns an item cold-start problem into a warm problem at one level up. It is the single most valuable idea in this case study, and the same trick appears in Section 4 (channel priors for new videos) and Section 9 (attribute-derived embeddings for new listings). Look for the entity one level above the cold item that is not cold.
Content features. Category and subcategory, title and description text embeddings, format (workshop, gig, market, class), indoor or outdoor, expected size, price, accessibility information, whether it is recurring.
Category priors. For each (user segment, category) pair, the historical registration and attendance rate. New event in the "live music" category, user segment "25–34, urban, previously attended 3 gigs" — the prior is a strong starting estimate before any event-specific signal exists.
Text similarity to the user's history. Embed the event description and compare against the average embedding of events the user has previously attended. Cheap, needs no interaction data for the new event, and it captures the flavour differences that categories miss — a "beginners' sourdough workshop" and an "advanced pastry masterclass" share a category and attract different people.
Recurring event lineage. Many events are instances of a series. The weekly running club has 50 previous instances with full history. Link instances to a series ID and inherit the series' statistics. A surprisingly large share of a typical events catalogue is recurring, which turns a large slice of the cold-start problem into a solved one.
User cold start
A new user has no history. The mitigations are cheaper here than in most problems:
- Ask. Onboarding asks for interests and a location. Users answer willingly because the connection to the product is obvious. Contrast with a video app, where "tell us what you like" feels like work.
- Location is available immediately and is the most powerful feature in the system, so a new user is much less cold than in other domains.
- Category priors by demographic fill the rest.
A brand-new Assembly user with a postcode and three declared interests can receive genuinely good recommendations on their first session. That is unusual and worth saying — the location constraint that makes item cold start worse makes user cold start better.
The inversion, stated
Standard recommendation design: build collaborative filtering as the core, add content features to patch cold start.
Event recommendation: build content and organiser features as the core, add collaborative signals opportunistically where they exist — for recurring series, for large events with real history, and for social signals from a user's connections.
Getting this inversion right is the difference between a system that works and one that recommends the same handful of large recurring events to everybody, which is what a collaborative-filtering-first design degenerates into here.
Features that matter here
With content first, the individual features deserve care. The whole system lives or dies on them.
Distance
The strongest single feature. Two decisions about it.
How to compute it. Straight-line distance is cheap and wrong in cities — a river or a railway can make 2 km straight-line into a 40-minute journey. Travel time by the user's likely mode is far better and costs an external call. The compromise used in practice: precompute travel-time estimates between geohash cells and look them up, which gives most of the accuracy at negligible request cost.
How to encode it. Not as raw kilometres. The relationship is strongly non-linear: 0–2 km is effortless, 2–10 km is a decision, 10–30 km is a commitment, beyond that it is a trip. Bucketed distance, or a log transform, both work. Trees handle this natively, which is one small argument in their favour for this problem.
Also encode distance relative to the user's own norm. Someone in central London who never travels beyond 4 km and someone in rural Devon who routinely drives 45 km have different distance functions. A feature of "distance divided by the user's median historical travel distance" captures that in one number.
Time until the event
Price
Price interacts with everything: with the attendance gap (paid registrations show up), with category norms (a £40 gig is normal, a £40 book club is not), and with the user's history. Encode as absolute price, as price relative to the category median, and as a free-or-paid flag. Price relative to category median is more informative than absolute price, and is the kind of feature that comes from thinking about the domain rather than from a checklist.
Category affinity
The user's historical registration and attendance rate per category, computed over trailing windows. Watch for two traps: a user with three registrations has affinity estimates that are mostly noise, so smooth toward the population prior in proportion to the amount of evidence; and affinity computed over all history misses that interests change, so use both a 90-day and an all-time window and let the model weigh them.
Social signals
If Assembly exposes connections, "how many of your connections are attending" is among the strongest features in the system — people go to events with other people. Also useful: whether anyone from a group the user belongs to is attending, and whether the organiser is someone they follow.
Two cautions. Social features are absent for users with no connections, so the model must handle missingness rather than treating zero as "nobody is going". And they amplify: a well-connected event gets recommended more, gets more attendees, and gets recommended more — the popularity loop from the video recommendation monitoring lesson in a social costume.
Batch or streaming, per feature
The decision every feature needs, and a good table to produce in an interview.
| Feature | Freshness needed | Computation |
|---|---|---|
| User category affinity (all-time) | Days | Nightly batch |
| User category affinity (90-day) | Hours | Hourly batch |
| Organiser history statistics | Days | Nightly batch |
| Event content embedding | Once, at creation | On write |
| Distance from user to event | Per request | Live, with a precomputed geohash lookup |
| Days until event | Per request | Live, trivially computed |
| Current registration count | Minutes | Streaming |
| Registration velocity (last 6 h) | Minutes | Streaming |
| Connections attending | Minutes | Streaming |
| Weather forecast for the event date | Hours | Scheduled fetch |
Registration velocity needs streaming because it is how the system detects an event taking off. An event that gained 200 registrations in the last six hours is a different proposition from one that gained 200 over three weeks, and by the time a nightly batch notices, the event may have happened.
The model: the problem's shape
With the features settled, choose the model by looking at what this problem actually is.
- Feature set: roughly 60–120 engineered tabular features — distances, counts, rates, categorical IDs of moderate cardinality, time buckets.
- Training data: on invented Assembly figures, roughly 40 million labelled (user, event) impressions per year, of which perhaps 2 million are positives.
- Candidate set at serving: ~600 per request, already filtered.
- Signal: concentrated in feature interactions — distance crossed with category, lead time crossed with day of week, price crossed with user history.
- No raw unstructured input that must be understood end to end, beyond a text embedding of the description that can be precomputed and fed in as a feature.
That is a description of the problem gradient-boosted decision trees were built for.
Why trees win here
They find interactions automatically. A tree that splits on category, then on days-until-event, then on distance has expressed a three-way interaction with no feature engineering. Getting a linear model to do the same requires constructing the crosses by hand.
They handle mixed types and skew natively. No scaling, no normalisation, no log transforms. Distance in kilometres, price in pounds, and a count of connections attending all coexist without preparation. Trees split on order, so monotonic transforms are irrelevant to them.
Missing values are handled by the algorithm. Weather is missing for indoor events; social signals are missing for unconnected users; organiser history is missing for first-time organisers. Standard boosting implementations learn a default direction per split for missing values, which is better than imputing a number that means something else.
They are strong at this data size. 40 million rows with 100 features is well within the range where boosting is at its best, and it is below the range where deep models start to pull ahead.
They train fast. Minutes to an hour on one machine. That matters more than it sounds: a model you can retrain in twenty minutes gets iterated on ten times a week, and a model that takes eight hours gets iterated on twice.
They are interpretable enough to debug. Feature importances and per-prediction attributions are available cheaply. When someone asks "why did we recommend a 40 km away pottery class to this user", there is an answer.
The full model ladder
State all of it, in order:
| Rung | Model | What it buys |
|---|---|---|
| 0 | Distance ascending, then attendee count | The bar. Genuinely decent in dense cities |
| 1 | Logistic regression on ~20 hand-crafted features | Personalisation, calibrated scores, trivially explainable |
| 2 | Gradient-boosted trees on the full feature set | Automatic interactions, missing-value handling. The recommendation |
| 3 | Trees plus embedding features from a separate model | Text and category semantics, without going fully neural |
| 4 | Deep model with embeddings | Only if multi-objective heads or learned text representations are needed |
Rung 3 is the pragmatic middle and worth naming: precompute a text embedding of the event description and a learned category embedding, reduce them to a handful of dimensions, and feed those as features to the tree model. You get semantic signal without abandoning the model class that suits the rest of the data.
When a neural model would become right
Be specific about the conditions, because "it depends" is not an answer:
- Very high-cardinality identifiers become central. If per-organiser or per-venue identity matters at a scale of hundreds of thousands, embedding tables handle it and trees do not.
- Several objectives need one shared representation. Predicting registration, attendance, and repeat-booking jointly with a shared trunk is a neural pattern (see the news feed model lesson).
- Raw text or images need end-to-end learning. If the event photo genuinely drives registration and a precomputed embedding is not capturing it.
- The data grows by an order of magnitude and the feature engineering stops keeping up.
- A retrieval stage becomes necessary. If the candidate set grew past a few thousand, you would need embeddings and a two-tower model — but the serving lesson shows why it will not.
None of these hold for Assembly today. Saying "not yet, and here is what would change my mind" is a stronger answer than either extreme.
Loss and training
Binary cross-entropy on the attendance label where available, registration otherwise, with a lower sample weight on registration-only rows. Class weighting for the roughly 1:20 imbalance.
Split by time — train on events that occurred before a cut-off, test on later ones — for the reasons in Step 5: training, with a specific twist: because labels arrive a median of 18 days after the recommendation, leave a gap between the training and test periods equal to the label horizon. Otherwise the training set contains events whose attendance was not yet known when the model would have been deployed, which is a subtle form of leakage.
Add a split by city as a second evaluation: hold out three cities entirely and check that the model generalises to a market it has not seen. Assembly enters new cities regularly, and a model that only works where it has history is a model that cannot support expansion.