Course Content
Generative AI System Design Interview
11 sections · 27 lessons
RAG: retrieval quality, grounded generation and evaluation
The index is built. Now a question arrives, and this lesson follows it through the online half of the system: retrieving the right passages, turning them into a grounded answer with citations, and measuring each half separately so that when an answer is wrong you know which half to fix.
Retrieval quality
Retrieval sets a hard ceiling on the whole system: the generator cannot use a passage it was never given. Three techniques raise that ceiling: hybrid retrieval, re-ranking and query rewriting.
Dense retrieval
Embed the question, find the nearest chunk vectors, return the top k. Because it works on meaning rather than words, it finds "How much does the business plan cost?" against a chunk that says "Business tier pricing is $75 annually" with no shared vocabulary at all.
Its blind spot is exactness. Embeddings compress meaning, and in doing so they blur precisely the things that must not be blurred: product codes, error numbers, version strings, surnames, and any rare token the embedding model has weak representations for. Ask for error E-40219 and dense retrieval will happily return chunks about errors in general.
Sparse retrieval
Classic keyword search over an inverted index, scored by a function such as BM25, which rewards matching rare terms and discounts common ones.
Its strength is exactness — E-40219 matches E-40219 and nothing else. Its blind spot is paraphrase: a question with none of the document's words scores zero, no matter how well it matches in meaning.
Hybrid, which beats both
The two failure modes are complementary, so run both and merge. The standard merge is reciprocal rank fusion: score each chunk by summing 1 / (k + rank) over the lists it appears in, with k a small constant. It needs no score calibration between systems, which is what makes it robust — comparing a cosine similarity against a BM25 score directly does not work.
Illustrative numbers on a technical documentation corpus: dense-only recall@50 of 84%, sparse-only 79%, hybrid 92%. That 8-point gain over the better single method is a ceiling raise on the entire system, and it is the cheapest one available.
Re-ranking
Retrieval optimises for speed over millions of chunks, so it uses a bi-encoder: question and chunk are embedded separately, which is what makes precomputation possible. The cost is that the two are never compared in detail.
A cross-encoder feeds the question and the chunk through a model together, so every word of the question can attend to every word of the chunk. It is substantially more accurate and cannot be precomputed — scoring the whole corpus per query would be impossible.
So use both: retrieve 50 candidates cheaply, re-rank them with the cross-encoder, keep the top 5. Fifty pairs cost around 40 ms and typically buy a large precision gain — an illustrative move from 61% to 82% precision@5. This is the best quality-per-millisecond purchase in the pipeline.
Query rewriting
The user asks "What does the business plan cost?" then "What about for a non-profit?"
The second question retrieves nothing useful. It contains no subject. Before embedding, rewrite it using the conversation history into a self-contained question: "What does the Acme business plan cost for a non-profit?"
A small model handles this in about 80 ms. Two refinements are worth knowing: query expansion generates several phrasings and retrieves for each, unioning the results, which raises recall on awkwardly-worded questions; and generating a hypothetical answer and embedding that instead of the question can help when questions and documents are written in very different registers, though it costs a generation and can drift.
Generation grounded in context
Retrieval has produced five passages. Now the generation half: how they are presented, how citations are produced, and how the system says "I do not know".
Prompt construction
Structure matters more than wording. A prompt that works:
You answer questions using only the sources provided below.Every factual claim must cite a source as [1], [2], and so on.If the sources do not contain the answer, say so and do not guess.SOURCES[1] (Acme Cloud › Pricing › Business tier, updated 2026-02-14) Business tier pricing is $75 annually per seat...[2] (Acme Cloud › Pricing › Non-profit discounts, updated 2026-01-30) Registered non-profits receive 40% off list pricing...QUESTIONWhat does the business plan cost for a non-profit?Four things are doing work. The sources are in a clearly delimited block — which is a safety control as well as a formatting one (see Safety, failure modes and follow-ups). Each source carries provenance and a date, so the model can prefer the more recent of two conflicting passages and the user can check. The citation format is specified rather than hoped for. And the abstention instruction comes before the sources, not after, because instructions at the very end of a long context compete with the retrieved text.
Ordering the passages
There is a well-documented effect in long-context generation: models use information at the beginning and end of a long context more reliably than information in the middle. The effect has been reproduced across model families and shows up strongly once the context is more than a few thousand tokens.
The practical consequence is free to act on. After re-ranking, place the highest-scoring passage first and the second-highest last, filling the middle with the rest. It costs nothing and measurably improves answers when you are supplying five or more passages.
The second consequence: do not pad the context. Supplying twenty passages instead of five because the window allows it makes answers worse, not better — it buries the good passage in the middle, and it costs four times the prefill.
Citations, and verifying them
Asking for citations gets you citations. It does not get you correct citations — models attach a plausible-looking marker to a claim the cited passage does not support.
So verify, cheaply and after the fact. For each cited claim, check that the cited chunk actually entails it, using a small natural-language-inference model or a cheap judge call at roughly 20 ms per claim. Claims that fail get their citation stripped and flagged, or the answer is regenerated. This is the difference between a system that cites and a system whose citations mean something.
Abstention, which is a feature
The instruction "if the sources do not contain the answer, say so" is the highest-value sentence in the prompt. A grounded system that answers everything is not grounded; it is a system that has learned to fill gaps from its own weights, which is exactly what you built retrieval to prevent.
Abstention needs to be designed as a product behaviour, not left as an edge case. A good abstention says what it could not find, shows the passages it did retrieve, and offers a next step — rephrase, or contact a human. A bad one is a bare "I don't know", which reads as broken.
Evaluating it needs two sets, and this is the part teams miss:
- An unanswerable set — questions whose answers are genuinely not in the corpus. Measure the abstention rate; you want it high.
- An answerable set — ordinary questions. Measure the false abstention rate; you want it low.
Optimising only the first gives a system that refuses everything, which is the chatbot's over-refusal failure from Section 4 wearing different clothes. Track both, always together.
Evaluating a RAG system
Two components, two evaluations, kept separate. This discipline is what lets you fix the right half, and it is the clearest single signal of experience in this case study.
Retrieval metrics
Retrieval is a ranking problem with ground truth, so ordinary information-retrieval metrics apply and there is nothing generative about it.
- Recall@k — of the chunks that could answer the question, what fraction appear in the top k? This is the ceiling on the entire system. If recall@5 is 60%, then 40% of questions are unanswerable no matter how good your generator is.
- Precision@k — of the k retrieved, what fraction are relevant? Low precision wastes context and, by the ordering effect described above, actively degrades answers.
- Mean reciprocal rank — how high up the first relevant chunk appears. Matters because of the ordering effect.
Measure recall at two values: recall@50 tells you whether your indexing and hybrid retrieval are sound, and recall@5 tells you whether your re-ranker is. A big gap between them is a re-ranker problem; a low recall@50 is an indexing or chunking problem.
Generation metrics
- Faithfulness (groundedness) — is every claim in the answer supported by the retrieved context? Measure by decomposing the answer into individual claims and checking each against the context with an entailment model or a judge. This is the metric that detects hallucination, and it is the one to report first.
- Answer relevance — does the answer address the question that was asked? A faithful answer can be faithful to the wrong passage.
- Citation accuracy — does each cited chunk support the claim attached to it? Distinct from faithfulness: an answer can be fully grounded and still cite the wrong source for each claim.
Diagnosing which half is broken
The 2×2 is the tool. Given a failing question, ask whether the correct chunk was retrieved and whether the answer was right:
| Answer correct | Answer wrong | |
|---|---|---|
| Correct chunk retrieved | Working | Generation problem — prompt, ordering, context too long, or the model ignoring context |
| Correct chunk not retrieved | Suspicious — the model answered from its own weights, which in a strictly-grounded system is a failure even though the answer was right | Retrieval problem — chunking, embedding, hybrid, or re-ranker |
The bottom-left cell is the one people miss and the most interesting. A correct answer with no supporting retrieval means the model is drawing on its own parametric knowledge. Today it is right; tomorrow, on a question where your corpus disagrees with the internet, it will be confidently wrong and there will be nothing in your metrics to warn you.
Building the evaluation set without a labelling budget
The bootstrap that makes this affordable: take a sample of chunks, and for each one use a model to generate a question that the chunk answers. You now have (question, known-relevant chunk) pairs — exact ground truth for every retrieval metric, at near-zero cost.
Have a human verify a sample of a few hundred, because generated questions are sometimes unanswerable or leak the answer's wording, which flatters retrieval. Then supplement with real user questions as they arrive, which are messier and more representative, and with a deliberately-built unanswerable set for abstention.