Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Similar listings: substitutability, baselines and incremental metrics


Prompt: "Design the 'similar listings' module on a vacation rental platform — the strip of recommendations shown on a listing page."

A focused case study on one technique: learning item embeddings from user sessions. It transfers to almost every recommendation problem in this course, and it is the cheapest large win available on most of them.

This lesson sets up the problem: what "similar" should mean here, the baselines that motivate embeddings, and the metrics — including the question of how many of the module's bookings would have happened anyway.

Sessions in, neighbours outClick sessionsSkip-gramover listingsListingembeddingsNearestneighboursThesimilar stripListings viewed together in one session are treated as context.
Similar means people considered them for the same trip, which captures neighbourhood and vibe that pixels and attributes miss.

The clarifying questions

  1. Similar by what? Price? Location? Architectural style? Or "people who considered this also considered that"? These give completely different results, and the fourth is the one that matches what users actually want.
  2. Where does the module appear, and what is it for? On a listing page while browsing, it is for exploration. On a listing page for a place that is unavailable on the user's dates, it is a rescue. On a booked-trip page, it is upsell. The rescue case is the highest-value one.
  3. What is it optimised for — clicks, bookings, or bookings that would not otherwise have happened? The metrics lesson argues for the third and admits how hard it is to measure.
  4. What is the inventory size and the market structure?
  5. What hard constraints apply? Availability on the user's dates, price range, guest count. These are filters, not preferences (see the event recommendation serving lesson).

The constraints we design against

Invented. The platform is Roost, a vacation rental marketplace.

ConstraintValue
Active listings~6 million
Markets (cities or regions)~500 meaningful ones
Listings per major market3,000 – 90,000
Search sessions per month~180 million
Median session length11 listing views
Sessions ending in a booking~4%
Module size12 listings, shown below the fold
Latency budget120 ms, since the module loads asynchronously

Framing it as machine learning

  • Input: an anchor listing, plus the user's context (dates, guests, price filters, session history).
  • Output: 12 listing IDs, ordered.
  • Objective: maximise the probability the user books one of them, conditional on not booking the anchor.

Task type: retrieval by embedding similarity, exactly as in Section 2 (Visual Search System) — with one crucial difference in where the embedding comes from.

Behaviour, not pixels

Section 2 learned an image embedding: two listings are similar if their photographs look alike. That is a genuine notion of similarity and it is the wrong one here.

Consider two Roost listings, both invented:

  • Listing A: a converted barn 8 km outside a small coastal town. Rustic photographs, exposed beams, £140 a night.
  • Listing B: a modern apartment in the same town centre. Bright, minimal, £135 a night.

Visually they share nothing. Structurally they share nothing. But people planning a weekend in that town look at both, and a traveller who cannot book A very often books B. Behaviourally they are close substitutes.

Now consider Listing C: a converted barn outside a different town, 300 km away. It looks almost identical to A. Nobody who wants A wants C, because location is the constraint that matters most and no visual model can encode "the same weekend trip".

So the definition of similarity that matters here is substitutability: two listings are similar if a traveller considering one would consider the other. That is a behavioural fact, not a visual or attribute one, and it is recorded in browsing sessions.

The baselines first

Three, in increasing order of quality, and all worth building:

  1. Same market, similar price, sorted by rating. Filter to listings within 20% of the anchor's price in the same market, sort by review score. No model. This is a real product and a fair chunk of the available value.
  2. Attribute similarity. Cosine similarity over an engineered attribute vector: price, bedrooms, property type, amenities, distance from the anchor, review score. Better, and it still cannot represent "these two feel like the same trip".
  3. Co-view counts. For each pair of listings, count sessions in which both were viewed, normalised by their individual popularity. Fully behavioural, needs no training, and it is the direct precursor of the embedding approach.

Co-view counts fail in a specific way that motivates the lesson on session embeddings: they are sparse. With 6 million listings there are 18 trillion possible pairs, and the great majority of genuinely similar pairs are never co-viewed by anyone. A listing with 40 views has co-view counts with perhaps 300 others, most of them 1. Embeddings solve exactly this: they generalise from observed co-occurrence to unobserved pairs, because similar listings end up in similar regions of the space even if no single session contains both.

Metrics

The offline metric here is unusually clean. The online metric hides a counterfactual question that most candidates never raise, and raising it is worth a great deal.

A clean offline metric and a hard online oneOffline is unusually clean• Held-out sessions give real next clicks• Rank the booked listing among negatives• Recall at k needs no annotationOnline hides a counterfactual• Bookings from the strip look like lift• The guest may have booked anyway• Only a holdout separates the two
Attributed bookings measure the module and the demand together, and raising that distinction is worth more than any number.

Offline metrics

The evaluation set builds itself from history. Take sessions that ended in a booking. For each one, pick a listing the user viewed before the booking as the anchor, hide the booked listing, and ask: does the model's similar-listings set for that anchor contain the listing they eventually booked?

MetricDefinition
recall@12Is the eventually-booked listing in the 12 shown? The headline
Average rank of the booked listingOver a larger candidate set, say the top 200 — a smoother signal than recall@12
recall@12, split by market sizeLarge and small markets fail differently
recall@12, split by anchor popularityPopular anchors have plenty of data; the tail does not

This metric is unusually good, for three reasons. The label is a booking — an expensive, deliberate action, not a click. It is available in volume without annotation. And it directly matches the product goal.

Its bias is the familiar one: it only contains listings that the search system surfaced. A perfect substitute the user was never shown cannot appear in the evaluation set, so the metric systematically favours models that resemble the current ranking.

Online metrics

MetricLayerReading
Module click-through rateImmediateIs the strip attractive?
Click-to-booking rate from the moduleSessionDo the clicks lead anywhere?
Bookings attributed to the moduleSessionThe headline
Bookings per session, platform-wideSessionThe number that actually matters — see below
Guard: bounce rate from the listing pageImmediateIs the module pulling people away from a listing they were about to book?

That guard metric matters more here than in most modules. A similar-listings strip sits on a page where the user is already considering a booking. Making it too compelling can cause deliberation rather than conversion — the user clicks away to compare, and books neither. Measuring only module-attributed bookings would score that as a success.

The counterfactual question

Here is the question worth raising unprompted, because it is the honest version of "did this work".

The module reports 40,000 bookings attributed to it last month. How many of those would have happened anyway?

Three possibilities behind that number:

  1. Genuinely incremental. The user was going to abandon — the anchor was unavailable, too expensive, or wrong — and the module rescued the session. Real value.
  2. Redirected. The user would have booked the anchor, and the module talked them into a different listing instead. Zero platform value, possibly negative if the substitute is cheaper or a worse match.
  3. Delayed but inevitable. The user would have found the same listing through search two minutes later. The module gets attributed credit for a booking that search would have produced.

Attribution counts all three the same way. Only the first is worth anything.

The way to measure it is a holdout experiment: turn the module off entirely for a small percentage of traffic and compare platform-wide bookings per session, not module-attributed bookings. If bookings per session are unchanged with the module off, the module is redirecting rather than creating.

An invented but instructive result: with the module on, 40,000 attributed bookings a month; in the holdout, total bookings per session fall by only 0.4%, implying perhaps 9,000 of the 40,000 were genuinely incremental. The module is still worth having — but it is worth about a quarter of what attribution claimed.