Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Smart Compose: the model, serving in 100 ms, and safety


The first half of Smart Compose fixed the frame: conditional generation of 3–10 tokens, 100 ms at the 90th percentile, two billion requests a day, trained under strict privacy controls. This lesson turns those constraints into a model, a training setup, a serving path and a safety layer — and shows how far arithmetic alone takes you.

The model, sized by the budget

Most design questions ask which model is best. This one asks which models are admissible, and the answer eliminates almost all of them before quality is discussed.

Which models are even admissibleDisqualified by arithmetic• Large decoder at 300 ms per call• Anything needing a cross-region hop• Beam search over a wide vocabularyAdmissible• Small LSTM or tiny transformer• Quantised and served near the user• Greedy decode over a short horizon
The question is not which model is best but which models fit inside the keystroke gap.

The disqualification, arithmetically

Take the 100 ms p90 budget and remove what is not negotiable:

StageBudget
Client → nearest serving location, round trip25 ms
Queueing and request handling8 ms
Safety filtering on the suggestion5 ms
Client render5 ms
Left for the model57 ms

Now the model has to do prefill over the context and decode 8 tokens inside 57 ms. From the inference cost and latency budget, a large served model decodes at roughly 20 ms per token at realistic batch sizes. Eight tokens is 160 ms of decode alone — nearly triple the entire remaining budget, before prefill.

A large model is not "expensive here". It is arithmetically impossible here. Say that in the interview and move on; do not spend three minutes weighing its quality.

The cost argument, which is independent and also decisive

Two billion requests a day. Suppose a large model costs $0.0003 per request (the worked figure from the cost and latency lesson). That is $600,000 a day, or over $200 million a year, for a convenience feature. Even if the latency worked, the economics would not.

What is admissible

A small decoder-only autoregressive model, in the region of a few hundred million parameters, with three compressions applied:

  • Distillation. Train the small model to imitate a large teacher's output distribution rather than only the raw text. This reliably beats training the same small model from scratch on the same data, because the teacher's full distribution carries more signal per example than a single correct token does.
  • Quantisation to 8-bit. Halves the bytes read per decode step and therefore roughly doubles the decode ceiling (see the cost and latency lesson), at a quality cost small enough to be worth it here. Measure it; do not assume it.
  • A short, hard-capped output. Cap generation at 10 tokens and stop early at a clause boundary. This is a modelling decision as much as a serving one — the model is trained on targets of that length, so it learns to produce complete short phrases rather than truncated long ones.

The quality cost, admitted honestly

A small distilled model is good at common phrasings and weak at unusual ones. It will complete Thanks for getting back to and I've attached the well. It will do poorly on a specialist sentence in an unfamiliar domain, and it will produce a bland suggestion where a larger model would have produced an apt one.

That trade is acceptable because of the failure cost. A bad suggestion costs the user one dismissive keystroke. Compare that with translation in Section 3, where a wrong output is published to a reader who cannot check it, or the retrieval system in Section 5, where a wrong output is a false statement of fact. The tolerable model size follows from the cost of being wrong, and naming that connection is what a strong answer sounds like.

Training and personalisation

Two questions here: what the model conditions on beyond the current sentence, and whether it adapts to the individual user.

Three places personalisation can liveGlobal model, all usersContext: thread, subjectPer-user adapter weightsRe-rank by past accepts
Cost and privacy exposure rise as personalisation moves deeper into the weights.

Conditioning on the email, not the sentence

Predicting from the current sentence alone throws away most of the available signal. Three context sources change suggestions materially:

  • The subject line. Re: Q3 budget review and Re: birthday drinks should not produce the same completion for I wanted to ask about.
  • The thread. The previous message sets topic, formality, and often the specific question being answered. A reply to a message ending in a question should be biased towards answering it.
  • The recipient. Writing to a manager and writing to a close colleague differ in register. The signal is the relationship, approximated by prior correspondence frequency and whether the address is internal — not by anything inferred about the person.

Architecturally, these are prepended to the model's input as a structured prefix:

Text
[SUBJECT] Q3 budget review[THREAD] ...last 200 tokens of the previous message...[RECIPIENT_TYPE] internal_frequent[BODY] Hi Devi, I wanted to ask about

Two things follow. The prefill cost grows with the prefix, which is exactly why prefill dominates in this system. And the prefix is nearly constant across the whole composition session, which sets up the single biggest serving optimisation in the serving design below.

Personalisation: three places it can live

Users have habits — a preferred sign-off, a way of opening, project names, colleagues' names. Capturing them raises acceptance materially. Where the personal model lives is the design decision.

On-deviceServer-side per-userNo personalisation
What it holdsA small adapter or phrase cache from this user's sent mailA per-user adapter or embeddingNothing
QualityGood on this device's habitsBest — sees all mail, all devicesBaseline
What it leaksNothing leaves the deviceA stored statistical model of one person's writingNothing
Cross-deviceDoes not transferConsistent everywheren/a
CostClient compute and batteryStorage and adapter-swap cost per userZero
DeletionDelete the app dataRequires a real deletion pathwayn/a

On-device is the privacy-preferable answer and the weaker product answer: it cannot see mail sent from another device, and phone compute limits how much adaptation is possible.

Server-side is better quality and creates a genuine obligation. A per-user adapter is personal data — it is a compressed model of how someone writes — so it needs retention limits, a deletion pathway that actually removes it, and access controls. It also creates a leakage route worth naming: a personalised suggestion can surface a phrase from one of the user's other emails into a draft they are composing over someone's shoulder.

The recommendation: start with no personalisation, prove the base model, then add on-device adaptation. Move to server-side only against a measured acceptance-rate gap, with deletion built before launch rather than after.

The training objective

Ordinary next-token cross-entropy over the body text, with the context prefix present but not scored — you want the model to use the subject line, not to learn to generate subject lines. Train on targets the length of real suggestions so the model learns to end cleanly at a phrase boundary rather than being truncated mid-word at the token cap.

Serving under a tight budget

Everything so far has been sizing. This is where the 100 ms budget is actually met, and four techniques do nearly all the work.

1. Trigger only at plausible points

The naive design sends a request on every keystroke. At two billion sessions' worth of keystrokes that is both unaffordable and pointless: mid-word contexts produce poor suggestions and the request is invalidated 150 ms later anyway.

A trigger policy runs client-side, costs nothing, and cuts request volume by a large factor:

  • Only at a word boundary — after a space or punctuation.
  • Only after at least 3 words have been typed in the current sentence.
  • Not inside a quoted reply, a signature block, a code block, or a pasted region.
  • Not within 200 ms of the previous trigger.
  • Not if the user dismissed the last two suggestions in this session — a strong signal they do not want them right now.

An illustrative outcome: from ~600 keystrokes per email down to ~20 trigger points. That is a 30× reduction in requests, achieved with client-side rules and no model.

2. Cache the prefix

From the training design above, the context prefix — subject, thread, recipient — is identical across every trigger in a composition session, and it is the bulk of the prefill cost. Compute its key-value cache once on the first request of the session, keep it keyed by session for a few minutes, and subsequent requests prefill only the new body tokens.

An illustrative split: a 250-token prefix and 30 body tokens becomes a 30-token prefill after the first request. Prefill drops by roughly 88% for 19 of the 20 requests in a session.

3. Stop early

Cap output at 10 tokens and stop at the first clause or sentence boundary. Half-suggestions are worse than short ones, and every token not generated is 4 ms returned to the budget.

4. Cancel in flight

When a new keystroke arrives, the in-flight request is predicting from a stale prefix. Cancel it — client-side, and propagate the cancellation to the server so the decode loop aborts and releases its batch slot. Without this, a fast typist generates a queue of obsolete work that delays the one request that still matters.

40Client debouncewait for a pause in typing25Network to edge15Context assemblyprior text, thread, locale30Prefillshort prompt, cached prefix60Decode 8 tokensthe only part that scales with output12Safety check25Network back207 ms total — the budget is about 300 ms before the suggestion stops feeling instantDecode is the largest single slice, and it is the one that grows if you let the model write longer suggestions.So the product decision — cap suggestions at a short phrase — is also the main latency decision. That coupling is what the interviewer wants you to name.
Seven stages share 207 ms, and only one of them gets longer when the model says more.

Cost per request, computed

Illustrative figures, using the $2.50 per GPU-hour from the cost and latency lesson, or $0.00069 per GPU-second, and a small quantised model at batch 64:

  • Prefill: 30 tokens ÷ 200,000 tokens/s aggregate = 0.00015 GPU-s
  • Decode: 8 tokens ÷ 40,000 tokens/s aggregate = 0.0002 GPU-s
  • Total: 0.00035 GPU-s × $0.00069 = $0.00000024 per suggestion

At 2 billion suggestions a day: about $480 a day, roughly $175,000 a year. Compare against the $600,000 a day a large model would have cost, and against the $14,400 a day this system would cost without the 30× trigger reduction.

Safety, monitoring and follow-ups

A suggestion is text the product puts into the user's mouth. That framing is the whole safety argument: the harm is not that the model said something, it is that the user's colleague will read it as something the user said.

Where a suggestion gets blockedCandidate textBlocked-termfilterSensitive-attributecheckConfidencethresholdShown to userSuppressing is always safe here, because showing nothing is a normal state of the product.
The harm is not that the model wrote it — it is that the recipient will read it as the user's own words.

What must be blocked, and how

Offensive or slur content. A small classifier on the generated phrase, at roughly 5 ms. Cheap, and it belongs in the 5 ms safety line of the stage budget above, not bolted on later.

Suggestions that assert a fact. The user types The deployment is scheduled for and the model offers Tuesday at 3pm. The model has no knowledge of the deployment. If accepted without care, the product has authored a false statement over the user's name. Mitigation: suppress completions that consist mainly of specific, checkable values — dates, times, quantities, names, identifiers — unless those values appear in the visible context. This is a constrained-decoding rule, not a classifier, and it is one of the higher-value controls here.

Words put in the user's mouth on sensitive topics. After I'm sorry to hear about your or Regarding your medical , the appropriate suggestion is none. Maintain a suppression list of sensitive contexts — bereavement, illness, dismissal, legal matters, finances — and do not trigger inside them. Users are not well served by autocomplete in a condolence email.

Assumptions about people. Google publicly removed gendered pronoun suggestions from Smart Compose in 2018 after finding it would suggest a gendered pronoun for an unnamed person, on the stated grounds that the cost of being wrong was high and the benefit small. That decision is the reference case for this class of problem: where a suggestion requires guessing an attribute of a person, do not suggest.

Multilingual support

Three genuine complications. Tokenisation efficiency differs sharply by script: text in languages that do not use spaces, or in non-Latin scripts, commonly costs substantially more tokens per character than English, which directly consumes the latency budget. Trigger rules are language-specific, since "after a word boundary" is not defined the same way everywhere. And a single multilingual model transfers knowledge to lower-resource languages at some cost to the highest-resource ones — the same trade-off the Google Translate model lesson works through in detail.

The practical answer: one multilingual model with per-language latency budgets and per-language trigger rules, and a per-language quality floor below which the feature is turned off rather than shipped badly.

Detecting drift from acceptance rate alone

Acceptance rate is the obvious health metric and it is heavily confounded. It moves with seasonality, with the mix of email types, with client releases that change how suggestions are rendered, and with the arrival of new users who have different habits.

A drop of 3% could be a model regression, or a UI change, or December.

Three things make it trustworthy:

  • A holdback. Keep 1% of users permanently on the previous model. Differences between the holdback and the treatment are attributable in a way absolute movements are not.
  • Cohort splits. Break acceptance rate by language, client, email length, and account age. A real regression usually concentrates in a segment; seasonality moves everything together.
  • A shadow evaluation set. The consented corpus from Data and privacy, scored on exact-prefix match after every model push, before any traffic moves.