Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Similar listings: learning embeddings from sessions


This is the core technique of this case study, and the one that transfers furthest. Instead of learning what listings look like, the model learns which listings people treat as alternatives, from nothing but the order they viewed them in.

The plain version of the technique works. Three modifications turn it from working into good, and each comes from a property of this domain that language does not have. This lesson covers both.

The analogy: a session is a sentence

Word embeddings are learned from the observation that words appearing in similar contexts have similar meanings. "Tea" and "coffee" both appear near "drink", "hot", "cup", "morning" — so a model trained to predict a word's neighbours ends up placing them close together, without anyone labelling them as related.

Apply the same idea to listings:

LanguageRoost
A sentenceA browsing session
A wordA listing viewed
Word orderThe order listings were viewed in
VocabularyThe 6 million listings
"Words in similar contexts have similar meanings""Listings viewed in similar sessions are substitutes"

A session like [barn_A → apartment_B → cottage_D → apartment_B → cottage_D → booked: cottage_D] is a sentence. Train a model to predict which listings appear near each other in such sentences, and listings that serve the same trip end up near each other in the embedding space.

Where the analogy breaks. Three places, and each has a design consequence:

  • Sessions are short. A median of 11 views against the thousands of words a language model sees per document. Less context per example, so you need many more sessions.
  • The vocabulary is enormous and churns. 6 million listings, with new ones appearing daily and old ones going inactive. Language vocabularies are around 100,000 and stable. This is why the cold start lesson exists.
  • There is an outcome. Sentences have no "correct" word, but sessions have a booking. That is extra supervision language models do not get, and Training details that decide quality, later in this lesson, shows how to exploit it.

Skip-gram, explained

Concretely, with a window of 3 and the session [A, B, C, D, E]:

  • Centre C produces the pairs (C, A), (C, B), (C, D), (C, E).
  • Centre B produces (B, A), (B, C), (B, D), and so on.

Each pair is a positive. Negatives are sampled listings not in the session (the training details below make this sampling much smarter).

Why co-occurrence beats attribute matching

Attribute matching answers "which listings have similar properties". Session co-occurrence answers "which listings do people treat as alternatives". These differ in ways that matter:

AnchorAttribute match saysSessions say
Barn, 8 km outside a coastal town, £140Other barns anywhereThe town-centre flat at £135, other places within reach of that beach
£600/night city penthouseOther £600 propertiesOther properties booked by people planning that kind of trip — which may include a £400 suite
Family house sleeping 8Other 8-sleeper housesTwo adjacent 4-sleeper flats, which no attribute model would ever surface

The third row is the one that makes the case. Nothing about a 4-bedroom house's attributes resembles two 2-bedroom flats, and yet families planning a group trip consider exactly that substitution. Behavioural embeddings capture it because people demonstrated it; attribute models cannot.

The embeddings also learn things nobody engineered: neighbourhood boundaries that do not match administrative ones, that certain listings serve business travellers and near-identical ones serve holidaymakers, and seasonal substitution patterns.

A browsing session is a sentence; each listing is a wordL-482L-119L-773L-205L-640sessionwindow of ±1predict theneighbourscontext: L-119context: L-205Listings that keep appearing near each other in real sessions end up near each other in the embedding space — without anyone labelling what "similar" means.the domain adaptation that mattersAdd the booked listing as a global context for every window in the session: the objective becomes "what did this browsingactually lead to", not just "what was clicked next".
Co-occurrence in a session is the label — which is why this works on click logs with no human annotation at all.

Training details that decide quality

The plain skip-gram model works. Three modifications turn it from working into good, and each comes from a property of this domain that language does not have.

One session, and the three modificationsviewedAviewedBviewedCviewedDbookedEnullcentreglobalcontextThe booking is context for every window, not only the last one.
Sessions that end in a booking, negatives drawn from the same market, and the booked listing as a global positive — each comes from a property language does not have.

1. The booked listing as a global positive

Sessions have an outcome. Use it.

In the standard model, a listing is a positive only for items inside its sliding window. But the booked listing is the answer to the whole session — it is what every listing viewed along the way was a step toward. So treat it as a positive for every window in the session, regardless of distance.

Concretely, in the session [A, B, C, D, E, F, G → booked: G], listing G is added as a positive for A, B, C, D, E, and F, not only for the items adjacent to it.

Why it works: it shapes the space around outcomes rather than around browsing paths. Two listings viewed near each other because a user was wandering are pulled together weakly by window co-occurrence. Two listings that both frequently precede the same booking are pulled together strongly. The second is a far better definition of similarity for a module whose purpose is to produce bookings.

The gain is largest at exactly the point of use: the module appears on a listing page and is trying to lead to a booking, which is the relationship this term explicitly encodes.

2. Negatives from the same market

This is the hard-negative idea from Step 3: data and labels in its most concrete form.

Sample negatives uniformly from all 6 million listings and almost every negative is in a different country. The model learns geography — a distinction the hard location filter already gives you for free — and stops there. Within a market, everything looks similar to it.

The fix: draw a substantial share of negatives from the anchor's own market. Now the model must distinguish a £140 barn outside a coastal town from a £180 cottage in the same town, which is the distinction the module actually needs.

A workable mix, and a good thing to state precisely:

Negative sourceShareWhat it teaches
Random from the full corpus~40%Coarse structure; keeps the space globally sensible
Random from the anchor's market~50%Fine-grained within-market distinctions — the important one
Listings shown in the session but not clicked~10%The user saw it and rejected it: the hardest and most informative negative

That third row is worth calling out. A listing that appeared in the search results, was seen, and was not clicked is a negative the user actually produced. It requires viewport logging, and it is the highest-quality negative available.

3. Sessions that end without a booking

96% of sessions end without a booking. Discarding them throws away most of the data; keeping them uncritically dilutes the signal.

The distinction to make is between two kinds of unbooked session:

  • Exploratory sessions — real browsing, listings compared, no decision reached. These carry genuine substitutability signal. Keep them, with the window-based positives, at a lower sample weight.
  • Noise sessions — one or two views, a bounce, a bot, an accidental tap. No signal. Filter them out: require a minimum of, say, 3 views and 30 seconds of activity.

So: booked sessions get full weight plus the global-positive term; qualifying exploratory sessions get partial weight and window positives only; short and abandoned sessions are dropped.

A refinement worth mentioning: sessions ending with a contact or enquiry to a host, or a saved listing, are intermediate — a real signal of intent short of a booking. Treat those as a weak version of the booked-listing term.

Other training decisions

Embedding dimension. 32 is a common and sensible choice for this problem. Item embeddings learned from co-occurrence need far fewer dimensions than image embeddings, because the underlying structure is lower-dimensional — location, price band, property style, trip type. At 6 million listings, 32 dimensions is about 768 MB in float32, which fits comfortably.

Window size. Around 5 works well. Sessions are short, so large windows include nearly the whole session and lose the ordering information; very small windows discard useful context.

Training frequency. Weekly full retrain. The vector space shifts on retraining, so the precomputed neighbour lists must be rebuilt at the same time and switched atomically — the same constraint as in the visual search monitoring lesson.

Aggregating to a user vector. Averaging the embeddings of a user's recent views produces a "where this session is heading" vector, usable to rank the module's candidates by session context rather than by the anchor alone. Cheap, and a good extension to propose.