Course Content
Generative AI System Design Interview
11 sections · 27 lessons
RAG: framing, why not fine-tuning, and the indexing pipeline
Design a system that answers questions over a body of documents the model was never trained on — a company's internal wiki, a product manual set, a legal archive — with citations back to the source.
This is the architecture most engineers reading this course will actually build, and it is the standard answer to the hallucination problem from What makes generative systems different. This section is written to stand alone: every term is defined where it appears, so it can be read before or after the others.
Clarifying questions
- What corpus, and how big? Ten thousand documents behaves very differently from ten million. Assume 500,000 documents averaging 3,000 words.
- What formats? Clean markdown is a weekend. Scanned PDFs with tables are a quarter.
- How often does it change? Hourly updates need an incremental indexing path; an annual archive does not.
- Are citations required? If yes, that changes the prompt, the chunking, and the evaluation. Assume yes — most useful deployments require them.
- Must answers be strictly grounded? May the model use its own general knowledge to fill a gap, or must every claim come from a retrieved document? These are different products. Assume strictly grounded.
- Who can see what? If different users may read different documents, access control is a retrieval-time concern and Safety, failure modes and follow-ups shows why it cannot be bolted on.
- Latency? Assume 2 seconds to first token.
- Are the questions self-contained? Conversational follow-ups like "what about the second one?" need query rewriting (see Retrieval quality).
The framing
Retrieval, then generation conditioned on what was retrieved. Two subsystems, designed separately, evaluated separately, and failing separately. This mirrors the retrieval-then- ranking split in the YouTube video search case study of Machine Learning System Design Interview, and the discipline is the same: keep them apart or you will not be able to tell which one is broken.
The two pipelines
Everything in this case study lives in one of two pipelines, and separating them is the first thing to draw:
Indexing (offline, runs when documents change): load → clean → chunk → embed → store, with metadata.
Query (online, runs per question): rewrite the question → embed → retrieve candidates → re-rank → assemble prompt → generate with citations → verify.
The offline pipeline decides your quality ceiling. The online pipeline decides your latency and cost. Candidates who describe only the online half have left out the part that determines whether the system works.
Why RAG rather than fine-tuning
"Why not just train the model on our documents?" is asked in nearly every interview on this topic and by nearly every stakeholder in real life. The answer is specific, and getting it right is a strong signal.
The one-line version
Fine-tuning teaches the model how to respond. Retrieval supplies what it should know.
Weights are a poor place to store facts. They are diffuse, unaddressable, unattributable, and expensive to change. An index is exactly the right place: addressable, attributable, cheap to update, and deletable.
The comparison in full
| Fine-tuning on the documents | Retrieval | |
|---|---|---|
| Adding a new document | Retrain — hours to days | Index it — seconds |
| A document changes | Retrain, or live with a stale model | Re-index that document |
| Removing a document on request | Not reliably possible without retraining | Delete the chunks |
| Citing a source | No mechanism — the model cannot say where a fact came from | Natural — you know which chunks you supplied |
| Corpus size | Limited by what training can absorb well | Limited by storage, so effectively unlimited |
| Per-user access control | One model per permission set — unmanageable | A filter on the retrieval query |
| Cost to set up | Hundreds to thousands of dollars | Embedding cost, roughly a few dollars per million chunks |
| Cost per request | Lower — short prompts | Higher — long prompts carrying retrieved context |
| Teaching output format | Excellent | Adequate via instructions and examples |
| Teaching domain style | Excellent | Weak |
Read the deletion row twice. If someone asks you to remove a document, retrieval lets you comply in seconds and fine-tuning does not let you comply at all without retraining. That alone settles the argument in most organisations.
When fine-tuning genuinely is the better answer
Four cases, and naming them is what stops this from sounding dogmatic:
- A consistent structured output at high volume. If every request must produce the same JSON shape and you are serving millions a day, a fine-tuned small model is cheaper per request and more reliable at format than a large model with a long instruction prompt.
- Domain style and vocabulary. Radiology reporting, legal drafting conventions, a specific house voice. These are not facts, and retrieval cannot supply them.
- A latency budget too tight for a retrieval hop. Retrieval adds 70–100 ms and a long prompt. Smart Compose (Section 2) cannot afford either.
- Behaviour, not knowledge. Refusal patterns, tone, and task decomposition are learned behaviours.
The combined answer, which is usually right
Fine-tune a smaller model for form — the output structure, the tone, the citation style, the refusal behaviour — and retrieve for facts. This is often cheaper and better than a large general model with a long prompt, because the fine-tune removes the need for two thousand tokens of instructions on every single request.
The indexing pipeline
The offline pipeline sets the ceiling on everything else. A chunk that was never indexed, or was indexed badly, cannot be retrieved by any amount of clever querying.
Step 1 — loading, which takes longer than you expect
Getting text out of real documents is the step that consumes most of the engineering time on real projects and gets one sentence in most interview answers. Say more than one sentence.
- HTML — strip navigation, footers, and cookie banners, or every chunk retrieves the same boilerplate.
- PDF — the hard case. Multi-column layouts read out of order without a layout-aware parser. Tables lose their structure and become word salad, which destroys any question about numbers. Scanned PDFs need optical character recognition, and its error rate propagates into every downstream stage.
- Spreadsheets and slides — meaning lives in structure and position, not in a text stream.
- Everything — preserve headings, section paths, and page numbers. You need them for citations and, as you will see below, for chunk quality.
Step 2 — the chunking decision
This is the single most consequential choice in a retrieval system and the one most often made by accepting a default.
Too small and context is lost. A 100-token chunk reading "The annual fee is $40 for the first year and $75 thereafter" is useless when the product name is in the heading two chunks above. It retrieves for the wrong queries and answers with an unattached number.
Too large and retrieval gets imprecise. A 2,000-token chunk covering six subtopics has an embedding that is the blur of all six, so it matches everything weakly and nothing strongly. It also consumes context you are paying for and, by the ordering effects in Generation grounded in context, buries the relevant sentence in the middle where the model attends least.
Four strategies, in increasing order of how much they help:
| Strategy | How | Best for |
|---|---|---|
| Fixed size | Every N tokens, with overlap | A baseline; unstructured text |
| Sentence-aware | Never split mid-sentence; pack to a target size | General prose |
| Structure-aware | Split on headings, keep sections intact | Documentation, manuals, wikis |
| Semantic | Split where consecutive sentence embeddings diverge | Long unstructured text with topic shifts |
The recommendation: structure-aware chunking targeting 300–600 tokens with 10–20% overlap between adjacent chunks, and — this is the highest-value trick in the pipeline — prepend the document title and heading path to every chunk before embedding it:
Acme Cloud Storage › Pricing › Business tierThe annual fee is $40 for the first year and $75 thereafter...That costs about fifteen tokens per chunk, requires no models, and repairs the most common retrieval failure there is: an orphaned passage that reads well and matches nothing.
Step 3 — embedding, and Step 4 — storage
Embed each chunk with an embedding model and store the vector alongside the chunk text and its metadata: source document, section path, page, last-modified date, document version, language, and — critically — the access-control identifiers that say who may read it (see Safety, failure modes and follow-ups).
Store the chunk text itself, not only the vector. You need it for the prompt, and re-fetching from the original document at query time adds latency and a failure mode.
Two practical notes. Embedding 500,000 documents at roughly eight chunks each is four million embedding calls — a few dollars at typical embedding prices and a few hours of wall-clock time, so plan it as a batch job. And changing the embedding model means re-embedding everything, because vectors from different models are not comparable. Version the index by embedding model and keep the old index serving while the new one builds.