Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Google Translate: framing, evaluation and parallel data
Design a machine translation service: text in one language goes in, the same meaning in another language comes out, at web scale.
Translation is the most measurable generative task in this course. There is no single correct output, but there are clearly wrong ones, and professional translators can be asked which of two versions is better. That makes it the right place to learn evaluation properly, which is what the metrics step of this lesson does.
Clarifying questions
- How many languages, and which directions? This is the first question because it changes the architecture completely. Ten languages is a different system from a hundred. Assume 100 languages for this case study.
- What content? Short user queries, whole documents, web pages, speech, or text inside photographs? Each has a different latency profile and a different notion of quality.
- Interactive or batch? A user waiting in a text box needs a response inside about 300 ms. A document translation job can take 30 seconds.
- Is formatting preserved? Translating a web page means keeping markup, links, and inline styling intact around moved words — a genuinely awkward engineering problem.
- What is the quality bar per language? Uniform quality across 100 languages is not achievable. Which pairs must be excellent, and which must merely be useful?
- Traffic. Assume 2 billion translation requests a day, heavily skewed towards short text and towards a small number of pairs. All figures in this case study are invented for teaching.
The framing
Sequence-to-sequence generation: a complete source sequence in, a complete target sequence out, where the output must preserve the meaning of the input and read naturally in the target.
Two words there are load-bearing. Complete — unlike Smart Compose in Section 2, nothing is being continued; the whole input is available before any output is produced. And preserve — this is a transformation with a constraint, not open generation. Adding a plausible detail is a failure here, whereas in a chatbot it might pass unnoticed.
Why encoder-decoder is the natural fit
The transformers lesson introduced the split. Here is where it pays off.
The source sentence is fixed and fully available, so it can be read bidirectionally — every word encoded with knowledge of the words after it as well as before it. That matters enormously for translation, because the correct rendering of a word frequently depends on what comes later in the sentence. In the bank was steep, bank is only disambiguated by steep.
The decoder then writes the target autoregressively while attending to that full source encoding. Source and target are separate objects with separate roles, and the architecture matches the task.
The two paths, decided now
Naming these early structures the rest of the answer:
| Interactive | Document | |
|---|---|---|
| Input | 1–3 sentences | Thousands of sentences |
| Latency budget | ~300 ms p90 | Seconds to minutes |
| Decoding | Beam search, narrow | Beam search, wider; more context per unit |
| Batching | Continuous, latency-capped | Large offline batches |
| Caching | Very effective — repeated strings | Ineffective — documents are unique |
Metrics, done properly
This is the part to study if you study one thing in this section. Translation is where the evaluation problem from Evaluating output with no correct answer becomes concrete enough to work through end to end.
BLEU: what it is and why it survives
BLEU compares a candidate translation against one or more human reference translations by measuring how many word n-grams they share — sequences of one, two, three, and four words — with a penalty for candidates that are too short.
Its weaknesses are documented and severe:
- It cannot see synonyms. The meeting starts at three and The meeting begins at three share almost no 4-grams. One is not worse than the other; BLEU says it is.
- It barely sees word order. Beyond 4-grams, any rearrangement is invisible.
- It is not comparable across languages or test sets. A BLEU of 35 on one pair and 28 on another says nothing about which system is better. This is the most common misuse.
- It is sensitive to tokenisation, so two teams can report different numbers for the same output. Standardised implementations exist precisely to remove this variable, and you should name one when asked.
- It correlates poorly with human judgement at high quality. When both systems are good, BLEU stops ranking them the way humans do — which is exactly when you most need it to.
And yet it stays, for a defensible reason: it is free, instant, reproducible, and directionally right on large changes. Its role is regression tripwire — fixed test set, fixed implementation, alert on a drop. That is not a quality measure, and describing it that way is the answer.
chrF, which does the same thing over character n-grams, behaves better for languages with rich morphology where whole-word overlap is rare. It costs nothing extra to report both.
Learned metrics
Newer metrics are neural models trained to predict human quality judgements from the source, the candidate, and usually a reference. They correlate substantially better with human ratings than BLEU, and they handle paraphrase, which is the main thing BLEU cannot do.
Three cautions to state alongside: they are models, so they have training distributions and fail outside them; they are harder to reproduce across versions; and because they are differentiable-ish targets, optimising directly against them invites the same gaming that optimising BLEU does. Use them as the primary automatic signal, pin the version, and keep BLEU as the cheap stable tripwire.
Human evaluation, which is the anchor
Two axes, and they fail independently:
- Adequacy — is the meaning of the source preserved? A fluent sentence that says something the source did not is an adequacy failure, and it is the dangerous kind.
- Fluency — does it read like natural target-language text? An accurate but stilted translation is a fluency failure, and it is the visible kind.
Collect it in one of three ways, in increasing cost and value: direct assessment (rate this translation 0–100), pairwise preference (which of these two is better — more reliable, as the evaluation lesson noted), and professional error annotation, where trained linguists mark each error with a category and a severity. Error annotation is the gold standard because it tells you what is wrong, not only that something is; it is also the slowest and most expensive, so it runs on hundreds of segments per release, not thousands.
Online signals
Users tell you things without being asked. Edit rate — how often a user modifies the returned translation — is the strongest single online signal. Copy rate suggests the output was useful. Re-translation into a third language suggests it was not. Explicit feedback exists and is sparse and skewed.
Data
Translation needs parallel data: the same content in two languages, aligned sentence by sentence. How much you can get differs by four orders of magnitude between pairs, and that single fact drives the model decision (see the multilingual model).
Where parallel data comes from
| Source | Scale | Quality |
|---|---|---|
| Official multilingual proceedings and publications | Millions of pairs, limited languages | Very high, but formal register only |
| Subtitles and dubbing scripts | Tens of millions | Colloquial and useful; timing-driven paraphrase |
| Localised websites and product documentation | Large and uneven | Good where professionally localised |
| Web-mined parallel text | Billions of candidate pairs | Noisy — this is where filtering earns its keep |
| Human-commissioned translation | Thousands | Highest, and by far the most expensive |
An illustrative shape: a high-resource pair might have 500 million usable sentence pairs. A low-resource pair might have 50,000. That is a factor of 10,000, and no amount of modelling cleverness closes it entirely.
Mining and filtering the web
The mining pipeline is worth being able to describe:
- Find candidate document pairs — same site, parallel URL structure, or matched by language-agnostic document embeddings.
- Align sentences within the pair — sequence alignment using length ratios and cross-lingual sentence embeddings, so sentence 7 in one document is matched to sentence 8 in the other where a sentence was merged.
- Filter aggressively. Discard pairs where the length ratio is implausible, where language identification disagrees with the claimed language, where the embedding similarity is below threshold, where either side is boilerplate, or where either side is itself machine translated — this last one matters and is easy to forget, because training on your own output degrades quality over time.
Filtering typically discards the large majority of mined candidates. That is the pipeline working, not failing.
The low-resource problem and back-translation
For a pair with 50,000 sentence pairs, a from-scratch model is poor. But monolingual text in the target language is usually abundant — news, books, web pages.
Back-translation turns that into training data:
- Train a rough model in the reverse direction, target → source.
- Run it over large amounts of real monolingual target-language text.
- You now have pairs whose source side is synthetic and whose target side is genuine human text.
- Add these to the training set for the forward direction.
The asymmetry is the whole trick. The model learns to produce the target, and the target side of every synthetic pair is real, well-formed human writing. Errors in the synthetic source side act more like noise than like corruption.
It works well and it is not magic: quality is capped by the reverse model, it can amplify that model's biases, and mixing too high a ratio of synthetic to genuine data degrades results. A common practice is to mix roughly one-to-one with real data and to tag synthetic examples so the model can distinguish them.
Pivoting is the other standard answer: translate a rare pair through a high-resource language. Cheap to build, and it compounds errors and loses nuance at each hop, so treat it as a fallback rather than a design.