Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Google Translate: the multilingual model, serving and safety


The first half of this case study settled the frame — a complete source in, a meaning-preserving target out, 100 languages, 2 billion requests a day — and the evaluation stack and data that will judge and feed the system. This lesson builds the rest: the model and how it is trained, the serving path with its budget and bill, and the failures a reader can never see.

The model

The architecture is an encoder-decoder transformer. The interesting decision is not that; it is whether you build one model per language pair or one model for all of them, and that decision has no clean answer.

Why per-pair is not an option at 100 languages

One hundred languages gives 100 × 99 = 9,900 directed pairs. Even at a modest model each, that is 9,900 training runs, 9,900 artifacts to store, version, and serve, and 9,900 quality regressions to monitor. And most of those pairs have almost no parallel data.

Per-pair models are correct for a small number of high-value pairs and impossible as a general strategy. Say the number out loud in the interview; it settles the question faster than an argument.

The multilingual model

One encoder-decoder trained on all pairs at once, with the target language indicated by a special token prepended to the input:

Text
[2fr] The meeting starts at three.   →   La réunion commence à trois heures.[2ja] The meeting starts at three.   →   会議は3時に始まります。

Three things this buys:

  • Transfer. Low-resource pairs borrow representations learned from high-resource ones, particularly within a language family. This is the largest single quality win available for the long tail.
  • Zero-shot directions. A model trained on English↔French and English↔Japanese can often translate French→Japanese without ever having seen that pair, because the encoder maps both into a shared representation. Quality is below a supervised pair, and it is far above nothing.
  • One artifact. One training pipeline, one deployment, one monitoring surface.

And one thing it costs, which you must name:

The middle grounds

Three ways to have some of both:

  • Scale the model. A larger multilingual model reduces the penalty because capacity is less contested. It costs more to serve, which at 2 billion requests a day is a serious constraint.
  • Language-family groupings. One model per family cluster — Romance, Slavic, Indic — keeps transfer where transfer helps most and limits interference between distant languages. Perhaps a dozen models rather than one or 9,900.
  • A shared model with per-language adapters. One base model, plus small parameter-efficient adapters per language or pair (see The build, fine-tune, or prompt decision), swapped at serving time. Recovers most of the high-resource quality at a small storage cost, and adds adapter-management complexity.

The recommendation: one large multilingual model as the base for all 100 languages, plus dedicated capacity — a bigger model or an adapter — for the small number of pairs that carry most of the traffic. It matches where the traffic and the quality expectations actually are.

Source textany of 100 languagesShared tokeniserone vocabularyEncoderlanguage-agnosticInterlinguashared spaceDecoderconditioned on target<2fr> target tagoutput in Frenchwhy one model, not 100 × 99Separate pairwise models would be 9,900 models. One multilingualmodel shares everything and needs only a target-language tag.zero-shot translationTrained on English↔French and English↔Tamil, the model can often translateFrench↔Tamil directly — because both were mapped into the same space.
The target-language tag is the whole trick: one decoder, steered rather than duplicated.

Training

Three things to cover: the objective, tokenisation across a hundred scripts, and the words that must not be translated.

The tokens that must survive intactMust not betranslatedNumbers and unitsPersonal namesURLs and emailsCode and markupUI placeholders
A subword vocabulary happily splits an address into pieces and reassembles it as nonsense.

The objective

Ordinary next-token cross-entropy on the target sequence, conditioned on the source. During training the decoder is fed the correct previous target tokens rather than its own predictions — teacher forcing — which makes training parallel across target positions and therefore fast.

Teacher forcing creates exposure bias: at training time the model always sees a perfect prefix, and at generation time it sees its own output, including its own mistakes. It has never been trained to recover from an error it made three tokens ago. In practice this matters less in translation than in longer-form generation, because outputs are short and the source keeps the decoder anchored — but it is the correct name for the phenomenon and the image captioning training lesson returns to it.

Label smoothing — training against a target distribution that puts a small amount of mass on tokens other than the correct one — is standard here, because it stops the model becoming over-confident and measurably improves the quality of beam search.

Tokenisation across scripts, which is harder than it looks

A multilingual model needs one shared vocabulary covering 100 languages. That vocabulary is built by a subword algorithm that learns frequent character sequences from the training corpus.

Three complications:

  • Languages without spaces. Chinese, Japanese, and Thai do not delimit words with spaces, so a whitespace-based pre-tokenisation step is meaningless. The standard fix is a subword algorithm that operates on raw text and treats whitespace as an ordinary character, so it learns segmentation from data rather than assuming it.
  • Vocabulary allocation. A shared vocabulary is allocated in proportion to training data. Under-represented languages get fewer dedicated subwords, so their text fragments into more tokens — commonly two to four times more tokens per character than English for some scripts. Every one of those tokens costs latency, costs money, and consumes context. Over-sampling low-resource languages when building the vocabulary is the standard mitigation, and it is worth naming because it is a cheap, high-value fix that teams forget.
  • Shared script, different language. Languages sharing a script share subwords, which helps transfer and occasionally causes the model to leak vocabulary between them.

Rare words, names, numbers, and the things that must not change

The dangerous failures in translation are not clumsy phrasings. They are a changed number, a dropped negation, or an invented name.

Four controls:

  • Copy behaviour. Subword tokenisation lets an unknown name be reproduced character by character, and models learn to copy rare source spans when nothing better is available. Test it explicitly.
  • Placeholder substitution. Before translating, replace numbers, dates, currencies, URLs, and code identifiers with typed placeholders; translate; substitute back. This makes it structurally impossible for £1,400 to become £1,000 — a correctness guarantee rather than a probability, which is why it is worth the pipeline complexity.
  • Do-not-translate markup. Honour explicit markers for product names and trademarks.
  • Terminology constraints. For enterprise use, force a glossary — a constrained decoding rule, not a prompt.

Serving

Two paths, one model family, and three optimisations that together decide whether this is affordable.

Beam search keeps two hypotheses aliveStartLeLachatcielchatte
Greedy commits to the first token forever; beam defers the commitment until the scores separate.

Decoding: beam search, and why

Decoding strategies said to use deterministic decoding where there is a best answer. Translation is the clearest case in the course: the user wants the correct rendering, not a creative one, and two users translating the same sentence should get the same result — which also makes caching possible.

Beam width 4 is a good default. Wider beams give diminishing returns and, past a point, actively degrade quality by favouring very short outputs, which is why the length penalty in beam search exists. Beam search costs roughly k times the decode compute, so width 4 means four times the per-token cost. That is a real bill, and it is the reason the document path can afford a wider beam and the interactive path cannot.

The three optimisations

1. Cache aggressively. Translation request traffic is enormously repetitive: interface strings, common phrases, popular queries, the same news headline translated by thousands of people. A simple key of (source text, source language, target language, model version) against a distributed cache is the single highest-value component in this system.

An illustrative hit rate of 60% on the interactive path is achievable and means 60% of requests never touch a GPU. Note the model version in the key — forgetting it is how a deployment gets silently reverted for two thirds of traffic.

2. Bucket batches by length. Batching sequences of very different lengths wastes compute on padding. Group requests into length buckets before batching and the waste largely disappears. This is worth roughly 20–30% of throughput on realistic traffic mixes.

3. Tier by request type. Not every request deserves the same model:

TierTraffic sharePath
Repeated short strings~60%Cache only
Interactive sentences~35%Mid-size multilingual model, beam 4, length-bucketed batches
Documents and low-resource pairs~5%Larger model or adapter, wider beam, offline batching

The latency budget, interactive path

Target 300 ms p90, keystroke-to-render:

StageBudget
Network round trip40 ms
Language detection5 ms
Cache lookup3 ms
Encoder pass over 25 source tokens15 ms
Decode 30 target tokens, beam 4180 ms
Detokenise, restore placeholders, safety check12 ms
Render10 ms
Total265 ms

Decode dominates, as it always does once the output is more than a handful of tokens — the opposite of Smart Compose in Section 2, where output was tiny and prefill dominated. That contrast is worth holding on to.

Cost per request, computed

Illustrative, at the $2.50 per GPU-hour from the cost and latency lesson ($0.00069 per GPU-second), mid-size encoder-decoder at batch 32:

  • Encoder: 25 tokens ÷ 60,000 tokens/s aggregate = 0.0004 GPU-s
  • Decode: 30 tokens × 4 beams = 120 token-steps ÷ 4,000 token-steps/s aggregate = 0.03 GPU-s
  • Total per uncached translation: 0.0304 × $0.00069 = $0.000021

Now apply the cache. At a 60% hit rate, the average cost per request is 0.4 × $0.000021 = $0.0000084. At 2 billion requests a day that is about $17,000 a day, roughly $6 million a year — against $42,000 a day, or $15 million a year, without the cache.

The cache is worth $9 million a year in this invented example. That is why it is the first thing to draw, not the last.

Safety, monitoring and follow-ups

Translation failures are unusual in one specific way: the person harmed usually cannot detect the failure, because if they could read the source they would not be using a translator.

The failures the reader cannot detectFailuresreaders missNegation droppedGender guessedFormality wrongSentence-only contextEntity mistranslated
If the user could read the source they would not need the translator, so no one reports these.

The failure that matters most: meaning inversion

A dropped negation turns do not take this medication with alcohol into take this medication with alcohol. A mishandled modal turns you may be eligible into you are eligible. These are fluent, confident, and catastrophic, and they are the canonical failure mode from What makes generative systems different.

Three mitigations, none complete:

  • Round-trip checking. Translate back to the source language and compare meaning against the original with a learned similarity metric. Cheap enough to run on high-stakes content, and it catches gross inversions well. It has false alarms, since a faithful translation may not round trip cleanly.
  • Negation and quantity consistency checks. Count negation markers, numbers, and named entities on both sides. A mismatch is a hard flag. This is a rule, not a model, and it catches the highest-severity errors.
  • Domain gating. For medical, legal, and safety-critical content, show a warning, offer a professional-translation path, and lower the threshold at which you decline to answer.

Gender, and why it is a design problem

Translating from a language without grammatical gender into one with it forces a choice the source did not make. The doctor said they would call into a gendered language requires picking a gender, and models trained on real text pick according to what was most common in that text — which encodes an occupational stereotype directly into the output.

Google has publicly documented shipping gender-specific translations for gender-neutral queries — returning both a feminine and a masculine version rather than silently choosing. That is the reference design for this class of problem, and it generalises: where the source is genuinely ambiguous, surface the ambiguity instead of resolving it invisibly.

Formality register

Many languages distinguish formal and informal address where English does not. Guessing wrong is a real social error — addressing a stranger informally, or a friend with excessive formality. Offer an explicit formality control, default it by content type (a business email formal, a chat message informal), and remember the setting for the session.

Context beyond the sentence

Sentence-by-sentence translation loses three things: pronoun antecedents from earlier sentences, consistent terminology across a document, and register consistency. The standard fix is a context window of the previous two or three source sentences (and, for terminology, a document-level glossary extracted on the first pass and applied on a second). Both cost context length and therefore money, which is why the document path can afford them and the interactive path mostly cannot.

Monitoring quality drift across 100 languages

The specific difficulty is scale: 9,900 directions cannot each have a human panel, and low-resource pairs degrade silently because almost nobody complains.

A workable scheme:

  • Automatic metrics per direction on a fixed test set, every build. Cheap, and catches large regressions.
  • Edit rate per direction, weekly, as a rate change rather than a level. Absolute edit rates are not comparable between languages; changes within a direction are.
  • Rotating human evaluation. Professional error annotation on a rotating subset, so every direction is reviewed on a schedule even if it cannot be reviewed continuously. Prioritise by traffic and by risk, not by traffic alone.
  • Round-trip similarity as a continuous background monitor on sampled live traffic.

The natural extensions

Speech translation as a cascade — recognise, translate, synthesise — where errors compound across three stages and latency is the sum of three budgets, versus end-to-end speech-to-text translation models which avoid the compounding and need scarcer training data. Image translation — detect text, recognise it, translate, and render it back into the image while matching the original layout, where the rendering is often the hardest part. Simultaneous translation, where the system must begin translating before the sentence is finished and therefore must decide when it has heard enough to commit — a genuinely different problem, since committing early risks a wrong word order and waiting costs latency.