Course Content
Generative AI System Design Interview
11 sections · 27 lessons
RAG: serving, safety and failure modes
Retrieval buys grounding and costs latency and money. This lesson prices both, then turns to the failure modes retrieval introduces and the controls they need.
The latency budget
Target: first token within 2 seconds.
| Stage | Latency | Notes |
|---|---|---|
| Query rewriting | 80 ms | Small model; skip when the question is self-contained |
| Query embedding | 10 ms | Cache by exact query text |
| Dense retrieval, top 50 | 20 ms | Approximate nearest neighbour over 4M vectors |
| Sparse retrieval, top 50 | 15 ms | Runs in parallel with dense |
| Fusion | 2 ms | |
| Cross-encoder re-rank, 50 → 5 | 40 ms | Batched in one pass |
| Prompt assembly | 5 ms | |
| Prefill, ~3,500 tokens | 290 ms | The dominant cost — this is the price of context |
| Time to first token | ~460 ms | Comfortably inside budget |
| Decode 300 tokens | 3,300 ms | Streams, so not user-blocking |
| Citation verification | 60 ms | Runs on the completed answer |
Two observations worth carrying. The whole retrieval stack is 167 ms — an eighth of the budget — while prefill over the retrieved context is 290 ms. The expensive part of retrieval is not retrieving; it is reading what you retrieved. And that is a direct argument against padding the context, on latency grounds as well as quality grounds.
Caching, at four levels
- Query embedding cache — exact-match on normalised query text. Cheap, and hit rates are higher than expected because popular questions repeat.
- Retrieval result cache — key on (rewritten query, access-control scope, index version). All three parts matter: omit the scope and you leak results across users; omit the index version and you serve stale chunks after a re-index.
- Full answer cache — key on the same plus the model version. Highest value, and only safe when the answer does not depend on anything user-specific beyond the access scope.
- Prefix cache on the static instruction block — the instructions and format specification are identical on every request, so order the prompt with them first (see Managing context) and the key-value cache is reused across all traffic.
Incremental index updates
A corpus that changes needs an update path that is not "rebuild everything".
- Upsert on change. Re-chunk the changed document, embed the new chunks, and replace that document's chunks atomically. Deleting old chunks before inserting new ones creates a window where the document is missing; insert then delete, keyed by document version.
- Delete on removal, including from any caches keyed on retrieval results.
- Track index version on every chunk, and include it in cache keys.
- Re-embedding is a migration. Changing the embedding model requires building a parallel index and switching over. Plan it as a blue-green deployment, not an update.
A useful practice: store a content hash per chunk and skip re-embedding chunks that did not change. A document edit usually touches one section, so this typically avoids most of the re-embedding work on an update.
Cost per request, computed
Illustrative, at $0.00069 per GPU-second:
- Query rewriting: 0.02 GPU-s → $0.000014
- Query embedding: 0.001 GPU-s → $0.0000007
- Retrieval: CPU-bound, roughly $0.000002 amortised over index infrastructure
- Cross-encoder re-rank, 50 pairs: 0.04 GPU-s → $0.000028
- Prefill, 3,500 tokens ÷ 12,000 tokens/s: 0.29 GPU-s → $0.00020
- Decode, 300 tokens ÷ 900 tokens/s: 0.33 GPU-s → $0.00023
- Citation verification: 0.03 GPU-s → $0.000021
- Total: about $0.00050 per question
Compare with an ungrounded answer — decode only, at about $0.00023. Grounding roughly doubles the cost per request, and about 40% of that increase is prefill over the retrieved context rather than the retrieval itself.
At 10 million questions a day: about $5,000 a day, or $1.8 million a year. Add index storage — four million vectors at 1,024 dimensions in 16-bit is roughly 8 GB, which is inexpensive — and the re-embedding cost whenever you change models.
Safety, failure modes and follow-ups
Retrieval reduces hallucination. It does not remove it, and it introduces three failure modes that did not exist before. Being able to name all four is what distinguishes a candidate who has run one of these systems.
Hallucination despite retrieval
The failure that surprises teams: the correct passage was retrieved, placed in the context, and the model still answered wrongly.
Four causes, each with a distinct fix:
- The model prefers its own knowledge. When retrieved context conflicts with what the model learned in training, it sometimes follows training. Fix: an explicit instruction that sources override prior knowledge, and measure faithfulness rather than assuming it.
- The context is contradictory. Two retrieved chunks disagree — an old policy and a new one. The model picks one silently. Fix: surface dates in the source block, instruct it to prefer recent sources and to say when sources conflict.
- The context is irrelevant but present. Retrieval always returns something. Given five irrelevant passages and no instruction to abstain, the model constructs an answer from them. Fix: a relevance threshold below which you retrieve nothing and abstain outright, plus the abstention instruction from Generation grounded in context.
- The answer needs synthesis across passages and the model gets the join wrong. Fix: multi-hop handling, below.
Retrieved content that is itself wrong
Your corpus contains a document that is out of date, or was wrong when written. The system will retrieve it, ground an answer in it, and cite it — which makes the wrong answer more credible than an ungrounded one would have been.
There is no model-side fix for this. The controls are corpus hygiene: recency metadata surfaced in the answer, an owner and review date per document, deprecation flags that down-rank or exclude, and a feedback path from users who spot an error back to the document owner. Treat corpus quality as an ongoing operational responsibility with a named owner, and say so in the interview — it is the answer that sounds like someone who has run this.
Prompt injection through retrieved documents
The chatbot safety lesson covered the mechanism. Retrieval is its main delivery route, because the attacker only needs to get text into a document your system will index — a support ticket, a wiki page, a shared file, a public web page.
The controls that matter here: a clearly delimited source block with an instruction that its contents are data (helps, does not solve), least privilege on any tools, no side-effecting actions without confirmation outside the model, and per-user access control so a poisoned document reaches only users who could already read it.
Multi-hop questions
"Which of our enterprise customers renewed after the 2025 price increase?" cannot be answered by one retrieval pass. No single chunk contains it. It requires finding the price-increase date, finding enterprise customers, finding renewal dates, and joining them.
Three approaches, honestly compared:
- Query decomposition. Break the question into sub-questions, retrieve for each, then synthesise. Works well for a small number of clean hops; costs one retrieval and one generation per hop.
- Iterative retrieval. Let the model retrieve, read, and decide whether it needs more — the tool loop from Tool use and agentic behaviour with retrieval as the tool. More flexible, harder to bound, and it needs the same iteration caps and budgets.
- Structured data. When the question is really a database query, route it to a database. Retrieval over prose is the wrong instrument for aggregation, counting, and joins, and recognising that is a strong answer rather than a concession.
Be honest that multi-hop is where retrieval-augmented generation is weakest. Each hop multiplies cost and compounds error, and a four-hop question with 85% per-hop accuracy is right about half the time.