Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

RAG: serving, safety and failure modes


Retrieval buys grounding and costs latency and money. This lesson prices both, then turns to the failure modes retrieval introduces and the controls they need.

Where the latency budget actually goes1540yes2570yes60120partly9002,400prefix onlyp50 msp95 msCacheableEmbed queryVector searchRe-rankGenerateRetrieval is a tenth of the budget; generation is the rest.
Optimising the vector index feels productive and moves the total by under five percent.

The latency budget

Target: first token within 2 seconds.

StageLatencyNotes
Query rewriting80 msSmall model; skip when the question is self-contained
Query embedding10 msCache by exact query text
Dense retrieval, top 5020 msApproximate nearest neighbour over 4M vectors
Sparse retrieval, top 5015 msRuns in parallel with dense
Fusion2 ms
Cross-encoder re-rank, 50 → 540 msBatched in one pass
Prompt assembly5 ms
Prefill, ~3,500 tokens290 msThe dominant cost — this is the price of context
Time to first token~460 msComfortably inside budget
Decode 300 tokens3,300 msStreams, so not user-blocking
Citation verification60 msRuns on the completed answer

Two observations worth carrying. The whole retrieval stack is 167 ms — an eighth of the budget — while prefill over the retrieved context is 290 ms. The expensive part of retrieval is not retrieving; it is reading what you retrieved. And that is a direct argument against padding the context, on latency grounds as well as quality grounds.

Caching, at four levels

  • Query embedding cache — exact-match on normalised query text. Cheap, and hit rates are higher than expected because popular questions repeat.
  • Retrieval result cache — key on (rewritten query, access-control scope, index version). All three parts matter: omit the scope and you leak results across users; omit the index version and you serve stale chunks after a re-index.
  • Full answer cache — key on the same plus the model version. Highest value, and only safe when the answer does not depend on anything user-specific beyond the access scope.
  • Prefix cache on the static instruction block — the instructions and format specification are identical on every request, so order the prompt with them first (see Managing context) and the key-value cache is reused across all traffic.

Incremental index updates

A corpus that changes needs an update path that is not "rebuild everything".

  • Upsert on change. Re-chunk the changed document, embed the new chunks, and replace that document's chunks atomically. Deleting old chunks before inserting new ones creates a window where the document is missing; insert then delete, keyed by document version.
  • Delete on removal, including from any caches keyed on retrieval results.
  • Track index version on every chunk, and include it in cache keys.
  • Re-embedding is a migration. Changing the embedding model requires building a parallel index and switching over. Plan it as a blue-green deployment, not an update.

A useful practice: store a content hash per chunk and skip re-embedding chunks that did not change. A document edit usually touches one section, so this typically avoids most of the re-embedding work on an update.

Cost per request, computed

Illustrative, at $0.00069 per GPU-second:

  • Query rewriting: 0.02 GPU-s → $0.000014
  • Query embedding: 0.001 GPU-s → $0.0000007
  • Retrieval: CPU-bound, roughly $0.000002 amortised over index infrastructure
  • Cross-encoder re-rank, 50 pairs: 0.04 GPU-s → $0.000028
  • Prefill, 3,500 tokens ÷ 12,000 tokens/s: 0.29 GPU-s → $0.00020
  • Decode, 300 tokens ÷ 900 tokens/s: 0.33 GPU-s → $0.00023
  • Citation verification: 0.03 GPU-s → $0.000021
  • Total: about $0.00050 per question

Compare with an ungrounded answer — decode only, at about $0.00023. Grounding roughly doubles the cost per request, and about 40% of that increase is prefill over the retrieved context rather than the retrieval itself.

At 10 million questions a day: about $5,000 a day, or $1.8 million a year. Add index storage — four million vectors at 1,024 dimensions in 16-bit is roughly 8 GB, which is inexpensive — and the re-embedding cost whenever you change models.

Safety, failure modes and follow-ups

Retrieval reduces hallucination. It does not remove it, and it introduces three failure modes that did not exist before. Being able to name all four is what distinguishes a candidate who has run one of these systems.

The four ways retrieval still failsRetrievalis not a fixHallucinates anywaySource itself is wrongInjection via docsMulti-hop questionsStale index
Grounding narrows the space of lies; it does not make the model unable to tell one.

Hallucination despite retrieval

The failure that surprises teams: the correct passage was retrieved, placed in the context, and the model still answered wrongly.

Four causes, each with a distinct fix:

  • The model prefers its own knowledge. When retrieved context conflicts with what the model learned in training, it sometimes follows training. Fix: an explicit instruction that sources override prior knowledge, and measure faithfulness rather than assuming it.
  • The context is contradictory. Two retrieved chunks disagree — an old policy and a new one. The model picks one silently. Fix: surface dates in the source block, instruct it to prefer recent sources and to say when sources conflict.
  • The context is irrelevant but present. Retrieval always returns something. Given five irrelevant passages and no instruction to abstain, the model constructs an answer from them. Fix: a relevance threshold below which you retrieve nothing and abstain outright, plus the abstention instruction from Generation grounded in context.
  • The answer needs synthesis across passages and the model gets the join wrong. Fix: multi-hop handling, below.

Retrieved content that is itself wrong

Your corpus contains a document that is out of date, or was wrong when written. The system will retrieve it, ground an answer in it, and cite it — which makes the wrong answer more credible than an ungrounded one would have been.

There is no model-side fix for this. The controls are corpus hygiene: recency metadata surfaced in the answer, an owner and review date per document, deprecation flags that down-rank or exclude, and a feedback path from users who spot an error back to the document owner. Treat corpus quality as an ongoing operational responsibility with a named owner, and say so in the interview — it is the answer that sounds like someone who has run this.

Prompt injection through retrieved documents

The chatbot safety lesson covered the mechanism. Retrieval is its main delivery route, because the attacker only needs to get text into a document your system will index — a support ticket, a wiki page, a shared file, a public web page.

The controls that matter here: a clearly delimited source block with an instruction that its contents are data (helps, does not solve), least privilege on any tools, no side-effecting actions without confirmation outside the model, and per-user access control so a poisoned document reaches only users who could already read it.

Multi-hop questions

"Which of our enterprise customers renewed after the 2025 price increase?" cannot be answered by one retrieval pass. No single chunk contains it. It requires finding the price-increase date, finding enterprise customers, finding renewal dates, and joining them.

Three approaches, honestly compared:

  • Query decomposition. Break the question into sub-questions, retrieve for each, then synthesise. Works well for a small number of clean hops; costs one retrieval and one generation per hop.
  • Iterative retrieval. Let the model retrieve, read, and decide whether it needs more — the tool loop from Tool use and agentic behaviour with retrieval as the tool. More flexible, harder to bound, and it needs the same iteration caps and budgets.
  • Structured data. When the question is really a database query, route it to a database. Retrieval over prose is the wrong instrument for aggregation, counting, and joins, and recognising that is a strong answer rather than a concession.

Be honest that multi-hop is where retrieval-augmented generation is weakest. Each hop multiplies cost and compounds error, and a four-hop question with 85% per-hop accuracy is right about half the time.