Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Smart Compose: framing, metrics and private data
Design inline sentence completion for an email client: as the user types, grey text appears ahead of the cursor suggesting how the sentence continues, and pressing Tab accepts it.
This is the first case study, and the one where a single constraint — latency — dictates every other decision. Work through it and you will have a template for reading a requirement as an architectural constraint rather than a nice-to-have.
Clarifying questions to ask first
Spend the first five minutes here. All figures used in this case study are invented for teaching.
- How much do we complete? The rest of the current word, the rest of the phrase, or the rest of the sentence? Longer completions are more valuable and much more likely to be wrong. Assume phrase-to-sentence length, around 3–10 tokens.
- How is a suggestion accepted? Tab or right-arrow to accept, any other key to dismiss. This matters because it defines the acceptance signal that becomes our main metric.
- When do we show one? Every keystroke, or only at plausible points? This turns out to be the biggest cost lever in the system.
- Does it personalise? Does it learn a specific user's phrasing, and if so where does that model live?
- What is the latency ceiling? Ask for a number and then sanity-check it against typing speed.
- Which languages, and is the same model used for all of them?
- What is the traffic? Assume 100 million composition sessions a day with roughly 20 trigger points each — two billion inference requests a day.
Why latency is the framing, not a requirement
A person typing at 40 words per minute produces a character roughly every 300 milliseconds. A fast typist at 80 words per minute produces one every 150 milliseconds.
A suggestion that arrives 400 milliseconds after the trigger appears after the user has already typed past the point it was predicting. It is not a slow suggestion; it is a wrong suggestion, because the context it was computed from is stale. Worse, it appears and then vanishes as the next keystroke invalidates it, which reads to the user as flicker.
So the useful budget is roughly 100 milliseconds end to end, measured at the 90th percentile, from keystroke to grey text on screen. Not an average — an average of 80 ms with a long tail means a visibly janky feature for the unlucky tenth.
The framing sentence
Conditional autoregressive text generation, with a hard 100 ms latency ceiling and a short output. Three parts, each doing work:
- Conditional — the generation is conditioned on the subject line, the thread, the recipient, and the text typed so far, not on the current sentence alone.
- Short output — 3–10 tokens. From the inference cost and latency budget, that inverts the usual cost picture: prefill will dominate decode, which is the opposite of a chatbot.
- Hard ceiling — the ceiling is not negotiable against quality. A better suggestion that arrives late is worth nothing, so the model is chosen by the budget rather than the budget being sized around the model.
Metrics
Smart Compose has an unusually clean online signal and unusually weak offline ones. Knowing which is which, and knowing the trap hiding in the clean one, is most of the metrics step.
Offline metrics, and why both are weak
Perplexity measures how surprised the model is by real text: lower means it assigned higher probability to what the human actually wrote. It is the standard language-model number and it is a poor proxy here for two reasons. It scores the model over all text, including the long stretches where we would never trigger a suggestion. And it rewards being unsurprised, which correlates only loosely with producing a suggestion someone wants to accept.
Next-token accuracy — how often the top prediction matches the next real token — is closer but still misaligned. Our unit of output is a phrase of 3–10 tokens, and a suggestion is accepted or rejected whole. Getting four of five tokens right is a rejected suggestion.
The better offline proxy is exact-prefix match at the suggestion level: given held-out real emails, trigger at plausible points, generate a suggestion, and check whether the user's actual continuation starts with it. Still a proxy, but it measures the thing the product does.
Online metrics, which are the real ones
| Metric | Definition | What it tells you |
|---|---|---|
| Acceptance rate | Accepted ÷ shown | Are the suggestions any good? |
| Trigger rate | Shown ÷ eligible points | Are we being too shy or too pushy? |
| Characters saved per email | Accepted characters ÷ emails | The actual product value |
| Post-accept deletion rate | Accepted then deleted within 5 s | Suggestions that looked right and were not |
| Composition time | Seconds from first keystroke to send | The guard metric |
Characters saved is the one to optimise. Acceptance rate is what a naive team optimises, and here is why that goes wrong.
The guard metric nobody sets up
The question the feature actually promises to answer is: do people finish emails faster?
That is measurable — time from first keystroke to send, per email, split by whether the user had the feature — but it is noisy, confounded by email length and by which users get which variant, and it needs a large experiment to detect a small effect. Run it anyway, at least once per major model change. It is the only metric that can catch the failure where every suggestion is accepted and users still take longer, because they are reading each suggestion, deciding, and being interrupted mid-thought.
Deletion-after-accept is the cheap early warning for the same failure: a rising rate means suggestions look right at a glance and are wrong on reflection.
Data and privacy
With the metrics fixed, the next question is what to train on. Training a model on email is a privacy problem before it is a machine-learning problem. This is a case where the constraint genuinely changes the architecture rather than adding a compliance checkbox, which is exactly why it is worth studying.
What makes this different from ordinary training data
Two properties compound. Email is highly personal — it contains medical details, financial details, relationship details, and identifiers for people who never agreed to anything. And generative models memorise and can reproduce their training data (see Data, licensing, and provenance), so a memorised sequence is not a theoretical exposure; it can be emitted as a suggestion into someone else's draft.
The failure to design against is concrete: a user types my account number is and the model helpfully completes it with a real number it saw during training. That is a reportable incident, not a quality bug.
The four controls, in order of how much they change the design
1. Nobody reads the data. Training runs against data no human inspects. This sounds like a policy statement and is actually an engineering constraint: it removes manual data inspection, error analysis on real examples, and eyeballing of failure cases — the debugging techniques teams rely on everywhere else. You have to replace them with aggregate statistics, synthetic probes, and a separately-collected consented evaluation set.
2. Aggressive filtering before anything else. Drop anything matching identifier patterns — numbers of certain shapes, email addresses, phone numbers, postal addresses, anything with high entropy. Drop rare sequences entirely: a phrase that appears in fewer than k distinct users' mail is exactly the phrase that would be memorised and is exactly the phrase nobody else needs suggested. This frequency threshold is the highest-value single control in the system.
3. Differential privacy during training. Add calibrated noise to gradients and clip each example's contribution, so the trained model is provably close to what it would have been without any particular user's data. The cost is real: noisy gradients mean lower quality at the same compute, and you pay for it in acceptance rate.
4. Federated learning. Train on devices, send only model updates, aggregate them centrally so no individual update is inspectable. Raw text never leaves the device. Google has published extensively on federated learning for mobile keyboard prediction, which is the closest well-documented analogue to this problem.
The evaluation problem this creates
If nobody can look at the training data, and production email cannot be exported, how do you run error analysis?
Three answers, all partial. Maintain a consented evaluation corpus — employees or paid participants who explicitly agree, giving you a few thousand real emails you may inspect. Use public correspondence corpora for regression testing, accepting the domain gap. And build a memorisation probe suite: a set of prompts designed to elicit identifiers, run before every release, checking that nothing plausible-looking and specific comes back.