Course Content
Machine Learning System Design Interview
11 sections · 33 lessons
Video search: framing, metrics and query-log data
Prompt: "Design video search. A user types a text query and gets back relevant videos."
Text in, videos out. The gap between those two things — a handful of words against an hour of audio and pixels — is what makes this case study the introduction to multimodal retrieval.
This lesson frames the problem, picks metrics that actually reveal when search is failing, and then looks hard at the data. Search has more usable signal than most problems in this course, and query logs are the richest data source here and the most systematically misleading.
The clarifying questions
- Do we search metadata only, or the video content itself? Title, description, and tags are text and cheap. Understanding what happens inside the video costs orders of magnitude more, and the visual content understanding lesson works out how much. Ask this first, because it sets the budget.
- Is search personalised? Two users typing "python tutorial" — should they see the same results? Personalisation helps engagement and hurts predictability, and it changes whether results can be cached.
- Multilingual? Does an English query match a Hindi video with English subtitles? This is a real product decision with a real cost.
- What is the catalogue size and the query volume?
- How fresh must results be? A video uploaded four minutes ago covering breaking news — should it be findable? "Yes" makes indexing latency a hard constraint.
The constraints we design against
Invented. The product is Nimbus, the video app from Section 1.
| Constraint | Value |
|---|---|
| Catalogue | 800 million videos |
| Upload rate | ~400 hours of video per minute |
| Search volume | 12,000 queries per second at peak |
| Latency budget | 300 ms p99 for the results page |
| Freshness | New uploads findable within 10 minutes |
| Query length | Median 3 words; 15% are a single word; 8% are a full sentence |
Framing it as machine learning
- Input: a text query, plus context (language, country, device, and optionally the user's history).
- Output: an ordered list of 20 video IDs.
- Objective: maximise the chance the user finds and watches something that satisfies the query.
The task type is retrieval, then ranking — and the most important framing decision is that these are two separate components with two separate metrics, built and evaluated independently.
| Retrieval | Ranking | |
|---|---|---|
| Input scale | 800,000,000 videos | ~500 candidates |
| Job | Do not lose the good video | Put the good video first |
| Metric | recall@500 | nDCG@20 |
| Budget | ~40 ms | ~70 ms |
| Feature richness | Few, precomputable | Many, including query-video crosses |
Trying to design one model that does both is the failure mode. A model rich enough to rank well cannot run over 800 million videos, and a model cheap enough to run over 800 million videos cannot rank well. This is the two-stage pattern from Step 6: serving, and this case study is where it earns its keep.
The baseline first
An inverted index over title, description, and tags with BM25 scoring, then sort by a blend of BM25 score and video view count. No machine learning at all beyond term statistics.
This baseline is genuinely strong, and saying so is a signal of judgement rather than a concession. For a query like "kerala backwaters drone footage", exact term matching against titles works extremely well. A large share of search traffic is navigational — people looking for a specific video or channel they already know about — and lexical matching handles that close to perfectly.
The baseline fails on vocabulary mismatch: a user searching "how to fix a leaky tap" will not match a video titled "Repairing a Dripping Faucet". Zero shared terms, perfect relevance. That specific failure is what the dense retrieval in the lexical and semantic retrieval lesson exists to solve, and framing it that way — as a named failure with a named fix — is much stronger than opening with "I'd use embeddings".
Metrics
Search has more usable signal than most problems in this course, and one metric that reveals failure better than any click-based number.
Offline metrics
Offline evaluation needs query–video pairs with relevance grades. Two sources:
Human relevance panels. Raters see a query and a video and grade relevance on a scale — say 0 (irrelevant) to 3 (perfect). Trained raters working from a written guideline. Expensive: at roughly 60 seconds per judgement and 20 videos per query, a 3,000-query panel is around 1,000 rater-hours. Slow, but it is the only source that can tell you about videos your system never showed.
Click logs as implicit relevance. Free, enormous, and biased in ways the data and features lesson covers in detail. Usable for iteration, not for absolute judgement.
The metrics themselves, split by stage:
| Stage | Metric | What it says |
|---|---|---|
| Retrieval | recall@500 | Is the relevant video in the candidate set at all? A miss here is unrecoverable |
| Retrieval | recall@500 by query type | Navigational, informational, and long-tail queries fail differently |
| Ranking | nDCG@20 | Graded relevance weighted by position — the headline ranking number |
| Ranking | MRR | Where the first relevant result lands; useful for navigational queries |
| Ranking | precision@1 | For queries with exactly one right answer |
Split every offline metric by query frequency band. Head queries (the top few thousand) behave completely differently from tail queries (asked once a month), and an aggregate number is dominated by the head while most of your quality problems live in the tail.
Online metrics
| Metric | Layer | Reading |
|---|---|---|
| Click-through rate on the results page | Immediate | Are results attractive? |
| Time to first click | Immediate | Rising means users are scanning further — the ordering is worse |
| Watch time from search-originated views | Session | Did the click lead to something worth watching? |
| Searches per session | Session | Ambiguous — could mean engagement or could mean failure |
| Search abandonment rate | Immediate | The share of searches with no click at all |
| Query reformulation rate | Immediate | The share of searches followed within 30 s by a modified query |
Abandonment and reformulation
These two deserve their own section because they are better failure detectors than click-through rate, and naming them is a strong signal.
Click-through rate is a rate over a numerator you control. Make thumbnails more sensational and it rises while relevance falls. It also cannot distinguish "no good results existed" from "good results existed and were ranked badly".
Abandonment — the user searched and clicked nothing — is unambiguous. Either nothing was relevant, or nothing looked relevant. Either way the search failed.
Reformulation is even sharper. A user typing "leaky tap", getting nothing useful, and retyping "dripping faucet repair" has told you precisely how your system failed and what the right answer looked like. Reformulation pairs are also, conveniently, one of the best sources of query-synonym training data available.
An invented but illustrative pattern: after a ranking change, click-through rate rises from 34% to 36% and looks like a win, while reformulation within 30 seconds rises from 11% to 15%. Users are clicking more and getting what they wanted less. Reformulation caught it; click-through rate hid it.
Data and features
Query logs are the richest data source in this course and the most systematically misleading. Understanding exactly how they mislead is most of the value in this step.
What the logs contain
Every search produces a row: the query text, the ordered list of results shown, which positions were viewed on screen, which were clicked, and how long the resulting watch lasted. At 12,000 queries per second that is around a billion searches a day.
The obvious use: treat a click as "relevant" and a non-click as "not relevant". That is wrong in three distinct ways.
Position bias
The effect is large. Click rates typically fall steeply with rank: an illustrative shape is position 1 receiving several times the clicks of position 5, with a long decay after. Exact numbers vary enormously by surface, query type, and device, and any figure quoted here would need measuring on your own logs — but the shape is consistent everywhere it has been studied.
Three corrections, in increasing order of effort:
- Model the position explicitly. Include position as a feature during training and set it to a constant at serving time. The model attributes part of the click to position rather than to relevance.
- Inverse propensity weighting. Estimate the probability that a result at each position is examined, then weight each training example by the inverse of that probability. A click at position 9 counts for much more than a click at position 1, because it survived a much lower chance of being seen.
- Randomisation experiments. On a small traffic slice, shuffle the top results. The resulting clicks give an unbiased estimate of the examination probability at each position. Costly in short-term quality and the only clean measurement available.
The other two biases
Presentation bias. A video with a bright, high-contrast thumbnail and a punchy title gets clicked more than an equally relevant video that looks dull. Train on clicks and you build a thumbnail-attractiveness model wearing a relevance model's name. Mitigate by using post-click signals — watch time, completion — rather than the click alone.
Selection bias. You only observe clicks on videos you showed. A perfect answer that the retrieval stage never surfaced generates no signal and can never be learned from. This is the feedback loop from Step 7: monitoring and the feedback loop in its search-specific form, and it is why the human relevance panel from the metrics lesson cannot be dropped.
Video-side signals
| Signal | Availability | Cost | Reliability |
|---|---|---|---|
| Title | Always | Free | Medium — creators optimise it |
| Description and tags | Usually | Free | Low — heavily spammed |
| Channel name and history | Always | Free | High |
| Automatic transcript | Requires speech recognition | High (see Visual content understanding) | High where speech is clear |
| On-screen text | Requires optical character recognition on frames | High | Medium |
| Thumbnail content | Requires an image encoder | Medium | Medium |
| Frame content | Requires a video encoder | Very high | Medium |
| Engagement history | After the video has been shown | Free | High, but zero for new uploads |
| Human-provided captions | Rare | Free | Very high |
Transcripts are the highest-value non-metadata signal by a wide margin, because most of the searchable meaning in most videos is spoken aloud. A cooking video's title says "Easy Weeknight Dinner"; the transcript says "chicken thighs", "harissa", "sheet pan", "twenty-five minutes". The transcript is what a user's query actually matches.
Query-side features
- Normalised query text — lowercased, with spelling corrected. Spelling correction alone materially improves tail-query recall.
- Query language, detected.
- Query intent class — navigational (looking for a specific channel or video), informational (how-to, explanation), or browse (entertainment, "funny cat videos"). Each wants a different result mix, and this is a small classifier trained on labelled queries.
- Query embedding, for the semantic retrieval path in the lexical and semantic retrieval lesson.
- Query frequency band, because head and tail need different treatment.