Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

RAG: framing, why not fine-tuning, and the indexing pipeline


Design a system that answers questions over a body of documents the model was never trained on — a company's internal wiki, a product manual set, a legal archive — with citations back to the source.

This is the architecture most engineers reading this course will actually build, and it is the standard answer to the hallucination problem from What makes generative systems different. This section is written to stand alone: every term is defined where it appears, so it can be read before or after the others.

The query pipeline, end to endUser questionEmbed andretrieveRe-rank to top 5Assemblethe promptGeneratewith citationsA second, offline pipeline chunks and indexes the documents before any of this runs.
Two pipelines share one index: ingestion writes it slowly, the query path reads it in milliseconds.

Clarifying questions

  • What corpus, and how big? Ten thousand documents behaves very differently from ten million. Assume 500,000 documents averaging 3,000 words.
  • What formats? Clean markdown is a weekend. Scanned PDFs with tables are a quarter.
  • How often does it change? Hourly updates need an incremental indexing path; an annual archive does not.
  • Are citations required? If yes, that changes the prompt, the chunking, and the evaluation. Assume yes — most useful deployments require them.
  • Must answers be strictly grounded? May the model use its own general knowledge to fill a gap, or must every claim come from a retrieved document? These are different products. Assume strictly grounded.
  • Who can see what? If different users may read different documents, access control is a retrieval-time concern and Safety, failure modes and follow-ups shows why it cannot be bolted on.
  • Latency? Assume 2 seconds to first token.
  • Are the questions self-contained? Conversational follow-ups like "what about the second one?" need query rewriting (see Retrieval quality).

The framing

Retrieval, then generation conditioned on what was retrieved. Two subsystems, designed separately, evaluated separately, and failing separately. This mirrors the retrieval-then- ranking split in the YouTube video search case study of Machine Learning System Design Interview, and the discipline is the same: keep them apart or you will not be able to tell which one is broken.

The two pipelines

Everything in this case study lives in one of two pipelines, and separating them is the first thing to draw:

Indexing (offline, runs when documents change): load → clean → chunk → embed → store, with metadata.

Query (online, runs per question): rewrite the question → embed → retrieve candidates → re-rank → assemble prompt → generate with citations → verify.

The offline pipeline decides your quality ceiling. The online pipeline decides your latency and cost. Candidates who describe only the online half have left out the part that determines whether the system works.

Why RAG rather than fine-tuning

"Why not just train the model on our documents?" is asked in nearly every interview on this topic and by nearly every stakeholder in real life. The answer is specific, and getting it right is a strong signal.

Teaching facts against teaching behaviourRetrieval fits facts• New documents arelive the moment indexed• Citations point at a real source• Access control can be enforced per userFine-tuning fits behaviour• Format, tone and domain vocabulary• Cuts prompt length, so cuts cost• Retraining needed for every fact change
Fine-tune for how the model should speak; retrieve for what it should know today.

The one-line version

Fine-tuning teaches the model how to respond. Retrieval supplies what it should know.

Weights are a poor place to store facts. They are diffuse, unaddressable, unattributable, and expensive to change. An index is exactly the right place: addressable, attributable, cheap to update, and deletable.

The comparison in full

Fine-tuning on the documentsRetrieval
Adding a new documentRetrain — hours to daysIndex it — seconds
A document changesRetrain, or live with a stale modelRe-index that document
Removing a document on requestNot reliably possible without retrainingDelete the chunks
Citing a sourceNo mechanism — the model cannot say where a fact came fromNatural — you know which chunks you supplied
Corpus sizeLimited by what training can absorb wellLimited by storage, so effectively unlimited
Per-user access controlOne model per permission set — unmanageableA filter on the retrieval query
Cost to set upHundreds to thousands of dollarsEmbedding cost, roughly a few dollars per million chunks
Cost per requestLower — short promptsHigher — long prompts carrying retrieved context
Teaching output formatExcellentAdequate via instructions and examples
Teaching domain styleExcellentWeak

Read the deletion row twice. If someone asks you to remove a document, retrieval lets you comply in seconds and fine-tuning does not let you comply at all without retraining. That alone settles the argument in most organisations.

When fine-tuning genuinely is the better answer

Four cases, and naming them is what stops this from sounding dogmatic:

  1. A consistent structured output at high volume. If every request must produce the same JSON shape and you are serving millions a day, a fine-tuned small model is cheaper per request and more reliable at format than a large model with a long instruction prompt.
  2. Domain style and vocabulary. Radiology reporting, legal drafting conventions, a specific house voice. These are not facts, and retrieval cannot supply them.
  3. A latency budget too tight for a retrieval hop. Retrieval adds 70–100 ms and a long prompt. Smart Compose (Section 2) cannot afford either.
  4. Behaviour, not knowledge. Refusal patterns, tone, and task decomposition are learned behaviours.

The combined answer, which is usually right

Fine-tune a smaller model for form — the output structure, the tone, the citation style, the refusal behaviour — and retrieve for facts. This is often cheaper and better than a large general model with a long prompt, because the fine-tune removes the need for two thousand tokens of instructions on every single request.

The indexing pipeline

The offline pipeline sets the ceiling on everything else. A chunk that was never indexed, or was indexed badly, cannot be retrieved by any amount of clever querying.

Step 1 — loading, which takes longer than you expect

Getting text out of real documents is the step that consumes most of the engineering time on real projects and gets one sentence in most interview answers. Say more than one sentence.

  • HTML — strip navigation, footers, and cookie banners, or every chunk retrieves the same boilerplate.
  • PDF — the hard case. Multi-column layouts read out of order without a layout-aware parser. Tables lose their structure and become word salad, which destroys any question about numbers. Scanned PDFs need optical character recognition, and its error rate propagates into every downstream stage.
  • Spreadsheets and slides — meaning lives in structure and position, not in a text stream.
  • Everything — preserve headings, section paths, and page numbers. You need them for citations and, as you will see below, for chunk quality.

Step 2 — the chunking decision

This is the single most consequential choice in a retrieval system and the one most often made by accepting a default.

Too small and context is lost. A 100-token chunk reading "The annual fee is $40 for the first year and $75 thereafter" is useless when the product name is in the heading two chunks above. It retrieves for the wrong queries and answers with an unattached number.

Too large and retrieval gets imprecise. A 2,000-token chunk covering six subtopics has an embedding that is the blur of all six, so it matches everything weakly and nothing strongly. It also consumes context you are paying for and, by the ordering effects in Generation grounded in context, buries the relevant sentence in the middle where the model attends least.

Four strategies, in increasing order of how much they help:

StrategyHowBest for
Fixed sizeEvery N tokens, with overlapA baseline; unstructured text
Sentence-awareNever split mid-sentence; pack to a target sizeGeneral prose
Structure-awareSplit on headings, keep sections intactDocumentation, manuals, wikis
SemanticSplit where consecutive sentence embeddings divergeLong unstructured text with topic shifts

The recommendation: structure-aware chunking targeting 300–600 tokens with 10–20% overlap between adjacent chunks, and — this is the highest-value trick in the pipeline — prepend the document title and heading path to every chunk before embedding it:

Text
Acme Cloud Storage › Pricing › Business tierThe annual fee is $40 for the first year and $75 thereafter...

That costs about fifteen tokens per chunk, requires no models, and repairs the most common retrieval failure there is: an orphaned passage that reads well and matches nothing.

Step 3 — embedding, and Step 4 — storage

Embed each chunk with an embedding model and store the vector alongside the chunk text and its metadata: source document, section path, page, last-modified date, document version, language, and — critically — the access-control identifiers that say who may read it (see Safety, failure modes and follow-ups).

Store the chunk text itself, not only the vector. You need it for the prompt, and re-fetching from the original document at query time adds latency and a failure mode.

Two practical notes. Embedding 500,000 documents at roughly eight chunks each is four million embedding calls — a few dollars at typical embedding prices and a few hours of wall-clock time, so plan it as a batch job. And changing the embedding model means re-embedding everything, because vectors from different models are not comparable. Version the index by embedding model and keep the old index serving while the new one builds.

Indexing — offline, run on a scheduleLoad documentsChunkEmbedStore vectors + metadatacost is paid once per document, not once per questionQuerying — online, per requestEmbed the questionRetrieve top-kRerankAssemble prompt + generatethe same embedding model, bothsidesthe answer is grounded in what was retrieved, and cites itMost RAG quality problems are created during indexing — bad parsing, bad chunking, the wrong embedding model — and only become visible at query time.
The two pipelines must share an embedding model — different models on each side produce silently wrong retrieval.