Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Evaluating RAG Performance
An engineering team stood up in a review and said their RAG system was 92% accurate. Someone asked what that meant. It turned out to mean this: on a set of 200 test questions, the correct source document appeared in the top 5 retrieved chunks 184 times.
Support had a different number. They had sampled 60 real conversations that week and judged 21 of the answers wrong or misleading — about 65% correct. Both numbers were honestly obtained. Both were correct. They were measuring different things, and neither told anybody what to fix.
When they finally instrumented it properly, the picture was unambiguous. Of the 184 questions where the right document was retrieved, the generator produced a correct answer only 138 times. End to end: 138 out of 200, or 69%. The retriever was fine. The generator was throwing away one answer in four that had been handed to it on a plate.
That is the entire case for structured evaluation. A single number cannot tell you which half of a two-stage system is broken, and RAG has at least two halves. You need numbers that decompose.
A framework that decomposes
Evaluate at four levels, each answering a question the level below cannot.
| Level | Question it answers | Needs labels? | Run it |
|---|---|---|---|
| 1. Retrieval | Did we find the right documents, and rank them well? | Yes — relevant doc IDs per question | Every commit |
| 2. Generation | Given those documents, was the answer grounded and on-point? | Partly — LLM judge plus a human-labelled calibration set | Every commit |
| 3. End-to-end | Would a user call this answer correct? | Yes — gold answers | Before every release |
| 4. Production | Is it still working on traffic we never anticipated? | No — implicit signals only | Continuously |
The relationship between levels 1, 2 and 3 is roughly multiplicative:
For the team above: 0.92×0.75=0.69. That equation is the most useful thing in this lesson, because it tells you where your headroom is. Pushing retrieval from 0.92 to 0.97 buys 0.97×0.75=0.73 — four points. Pushing generation from 0.75 to 0.90 buys 0.92×0.90=0.83 — fourteen points, for less work. Teams spend their time on retrieval because retrieval is the fun part.
Measure both factors separately, then work on the smaller one. A system with 0.95 retrieval and 0.60 generation is not a retrieval problem, however much it feels like one.
Level 1: retrieval evaluation
You need a labelled set: for each question, the IDs of the documents that genuinely contain the answer. Two hundred questions is enough to start. Building it is a day of tedious work and it is the highest-return day you will spend on the project.
Take one question with three relevant documents in the corpus — d17, d42, d88 — landing at ranks 1, 3 and 12.
Precision@K and Recall@K
| K | Relevant found | Precision@K | Recall@K |
|---|---|---|---|
| 1 | 1 (d17) | 1/1 = 1.000 | 1/3 = 0.333 |
| 3 | 2 (d17, d42) | 2/3 = 0.667 | 2/3 = 0.667 |
| 5 | 2 | 2/5 = 0.400 | 2/3 = 0.667 |
| 10 | 2 | 2/10 = 0.200 | 2/3 = 0.667 |
| 20 | 3 (all) | 3/20 = 0.150 | 3/3 = 1.000 |
Read the two columns against each other. Precision falls monotonically as K grows; recall rises. There is no K that is good at both, so you have to decide which failure you can live with.
For RAG, recall matters more, and for a specific reason: a document you did not retrieve is unrecoverable — nothing downstream can invent it — whereas an irrelevant document is merely noise a re-ranker or a well-constrained generator can survive.
But precision is not free either, and this is where people over-correct. Push K to 20 and you are sending 20 chunks to the generator. Models measurably degrade when the relevant fact is buried in the middle of a long context, and irrelevant chunks raise hallucination rates. The standard resolution is to retrieve wide and re-rank narrow: high recall at K=50, high precision at K=5.
MRR — when there is one right answer
Take the reciprocal of the rank of the first relevant document. Across four questions with first-relevant ranks 1, 4, 2 and "not found in top 10", the reciprocal ranks are 1.000, 0.250, 0.500 and 0.000, giving MRR=1.75/4=0.4375.
MRR is the right metric for factoid lookup, where one document answers the question and you only care how near the top it lands. It is the wrong metric when several documents each contribute part of the answer, because it ignores everything after the first hit.
NDCG@K — when relevance comes in grades
NDCG handles graded relevance (3 = directly answers, 2 = useful context, 1 = tangential, 0 = irrelevant) and discounts each hit by its position.
Grades [2, 0, 3, 1, 0] at ranks 1-5 give
DCG@5=1.00002+0+2.00003+2.32191+0=2.0000+1.5000+0.4307=3.9307
The ideal ordering of those same grades is [3, 2, 1, 0, 0], so
IDCG@5=3.0000+1.58502+2.00001=3.0000+1.2619+0.5000=4.7619
and NDCG@5=3.9307/4.7619=0.825. Two warnings: an exponential-gain variant using 2reli−1 is equally common and gives different numbers, so always state which you used; and grading relevance on a four-point scale needs an annotation guideline, or two people will disagree on half the labels.
MAP@K — precision averaged over every hit
Average Precision takes Precision@i at each rank i where a relevant document appears, and averages. For our three documents at ranks 1, 3 and 12:
| Relevant doc at rank | Precision at that rank | Value |
|---|---|---|
| 1 | 1/1 | 1.000 |
| 3 | 2/3 | 0.667 |
| 12 | 3/12 | 0.250 |
AP=(1.000+0.667+0.250)/3=1.917/3=0.639. For a second question with relevant documents at ranks 2 and 5, AP=(1/2+2/5)/2=0.900/2=0.450. Across the two, MAP=(0.639+0.450)/2=0.544.
MAP is the metric to report when multiple documents matter and you want one number that respects both how many you found and where you put them.
| Use | When | Blind to |
|---|---|---|
| Recall@K | Choosing your fetch depth before re-ranking | Order entirely |
| Precision@K | Setting the final K sent to the generator | What you missed |
| MRR | Single-answer factoid lookup | Everything after the first hit |
| NDCG@K | Graded relevance; comparing re-rankers | Nothing much — but needs graded labels |
| MAP@K | Multiple relevant docs, binary labels | Degrees of relevance |
Level 2: generation evaluation
Now hold retrieval fixed. Feed the generator the correct context and ask what it did with it.
Faithfulness — is every claim supported by the context?
This is the metric that catches hallucination, and the method is to decompose the answer into atomic claims and check each one against the retrieved text.
Take a real answer:
"Enterprise annual contracts have a 14-day refund window from the invoice date. Refunds require approval from two signatories. Partial refunds are prorated monthly. The policy was last updated in March 2024."
Split it and check:
| # | Atomic claim | In context? |
|---|---|---|
| 1 | Enterprise annual contracts have a 14-day refund window | Yes |
| 2 | The window runs from the invoice date | Yes |
| 3 | Refunds require two signatories | Yes |
| 4 | Partial refunds are prorated monthly | No |
| 5 | The policy was last updated in March 2024 | No |
Faithfulness =3/5=0.60. Look closely at claims 4 and 5. They are not wild inventions — they are exactly the kind of plausible administrative detail a model adds because policy documents usually contain such things. That is what real hallucination looks like in production: not nonsense, but confident specifics in the same register as the true parts, which is precisely why humans reading the answer do not notice.
1FAITHFULNESS = """You are checking whether a claim is supported by context.23Context:4{context}56Claim: {claim}78Answer with exactly one word: SUPPORTED, CONTRADICTED, or NOT_STATED.9A claim is SUPPORTED only if the context states it or entails it directly.10Plausible inference is NOT support."""1112def faithfulness(answer, context, judge, splitter):13 claims = splitter(answer) # one factual assertion each14 verdicts = [judge.invoke(FAITHFULNESS.format(15 context=context, claim=c)).content.strip() for c in claims]16 return sum(v == "SUPPORTED" for v in verdicts) / len(claims)The last line of that prompt is doing real work. Without it, judges mark anything reasonable as supported, and faithfulness scores cluster around 0.95 regardless of how the system is behaving.
Calibrate the judge, or the judge is decoration
An LLM judge is a model, and unvalidated models lie. Label 100 claim-context pairs by hand, run the judge over the same 100, and compute agreement. Raw agreement is misleading when one class dominates, so use Cohen's kappa:
Suppose the judge and the human agree on 88 of 100 pairs, so po=0.88. The human marked 70 supported, the judge 76, so chance agreement is pe=(0.70)(0.76)+(0.30)(0.24)=0.532+0.072=0.604. Then κ=(0.88−0.604)/(1−0.604)=0.276/0.396=0.697.
Kappa around 0.70 is substantial agreement and usable for tracking relative change. Below about 0.60, your judge is not measuring what you think, and every conclusion you draw from it is noise. Re-check kappa whenever you change the judge model or the prompt.
Answer relevance — does it address the question asked?
An answer can be perfectly faithful and completely beside the point. The standard technique runs backwards: ask a model to generate the questions this answer would be a good response to, embed them, and compare each to the real question.
For the question "What is the refund window for enterprise annual contracts?", the generated questions and their cosine similarities might be 0.91, 0.88 and 0.62, giving relevance =(0.91+0.88+0.62)/3=2.41/3=0.803. The 0.62 is the informative one — it usually corresponds to a paragraph where the answer drifted into a related but unasked topic, which is a retrieval precision problem showing up as a generation symptom.
Fluency, conciseness and context utilisation
Modern models are fluent by default, so fluency is rarely your bottleneck and rarely worth measuring. Two cheaper proxies earn their place:
- Length ratio — answer tokens divided by gold answer tokens. Consistently above about 2.5 means the model is padding, restating the question, and hedging, all of which users read as evasive.
- Context utilisation — the fraction of retrieved chunks actually cited in the answer. Send 5 chunks, cite 1, and you are paying five times over for one document's worth of value. Utilisation below 0.4 is a signal to reduce K, not to improve the prompt.
Level 3: end-to-end evaluation
Levels 1 and 2 diagnose. Level 3 decides whether to ship. It needs gold answers, and the composition of that set matters more than its size.
| Slice | Share | Why it must be there |
|---|---|---|
| Single-fact lookup | 40% | The bulk of real traffic |
| Multi-document synthesis | 20% | Catches the retriever returning one facet of a many-part answer |
| Comparison and conditionals | 15% | Where negation and qualifiers break both stages |
| Ambiguous or under-specified | 10% | Tests whether the system asks rather than guesses |
| Unanswerable | 15% | The only way to measure whether it refuses |
That last row is the one everybody omits, and omitting it is the reason so many systems score well in evaluation and hallucinate in production. If every question in your test set has an answer in the corpus, a system that never refuses scores perfectly — and you have accidentally optimised for confident guessing. Include questions you know the corpus cannot answer and score the refusal as the correct response.
A/B testing two configurations, honestly
Run both configurations over the same questions and compare per question, not in aggregate. Suppose Config A gets 138 of 200 right and Config B gets 152 — a seven-point improvement. Ship it?
Not yet. Look at the paired outcomes:
| B correct | B wrong | |
|---|---|---|
| A correct | 114 | 24 |
| A wrong | 38 | 24 |
The 114 and the 24 in the corner tell you nothing — both configurations behaved identically there. The evidence lives entirely in the discordant cells: B fixed 38 questions A got wrong, and broke 24 that A got right. McNemar's test uses exactly those two numbers:
The critical value for one degree of freedom at the 5% level is 3.841. Since 2.726<3.841 (p ≈ 0.10), this result is not statistically significant. A seven-point improvement on 200 questions is well inside what you would expect from shuffling which questions happen to be in your test set.
Twenty-four questions got worse. If you ship on the aggregate number alone you will never look at them, and one of them will be the query your largest customer runs every morning.
If the same effect held at 500 questions — 95 fixed, 60 broken — then χ2=(35−1)2/155=1156/155=7.46, comfortably above 3.841. Either collect more questions or accept that you cannot yet tell the two configurations apart. And read the 24 regressions individually: they usually cluster into one recognisable pattern, and that pattern is more useful than the aggregate ever was.
Level 4: production monitoring
Offline sets go stale within weeks. Real users ask things nobody anticipated, in phrasing nobody wrote down. The trick in production is that you have no labels — so you monitor signals that correlate with failure without needing them.
| Signal | What a change means | Act when |
|---|---|---|
| Median retrieval top score | Corpus drifting away from query distribution | Drops more than 0.05 week on week |
| Score margin (1st − 5th) | Flat profiles mean the corpus has nothing specific | Median below 0.15 |
| Refusal rate | Rising: coverage gaps. Falling sharply: guardrail broke | Moves more than 5 points either way |
| Rephrase rate | User re-asks the same thing within 60 s — the answer failed | Above baseline + 5 points |
| Citation rate | Falling means answers coming from parameters, not context | Below 0.85 |
| Escalation to human | The most honest quality signal you have | Any sustained rise |
| p95 latency | Users abandon before quality matters | Above your stated budget |
Rephrase rate deserves the emphasis. If a user asks a semantically similar question within a minute of the last one, the first answer almost certainly failed them. It needs no labelling, no judge and no annotation budget — just an embedding comparison between consecutive queries in a session. One team's baseline sat at 8%; a deploy pushed it to 19% overnight, and the alert fired eleven hours before the first support ticket arrived.
Sample and label a hundred real conversations a week regardless. Automated signals tell you that something changed; only reading transcripts tells you what.
Reading the numbers back to a fix
The point of measuring in layers is that each combination of results points at one stage.
| Recall@50 | Recall@5 | Faithfulness | Diagnosis | Fix |
|---|---|---|---|---|
| Low (<0.85) | Low | — | The answer is not in the candidate set at all | Chunking, embedding model, add BM25 — a re-ranker cannot help |
| High | Low | — | Found but badly ordered | Add or upgrade the cross-encoder re-ranker |
| High | High | Low | Right context, ungrounded answer | Constrain the prompt to context, temperature to 0, require citations |
| High | High | High, relevance low | Right topic, wrong facet | Query condensation and rewriting; check chunk boundaries |
| High | High | High | Metrics good, users unhappy | Format, tone, latency, missing citations — measure the right thing |
Two failure modes deserve naming because they waste the most time.
Optimising the wrong factor. The team in the opening spent six weeks on embeddings and re-ranking, moving retrieval from 0.89 to 0.92. End-to-end went from 0.67 to 0.69. Two afternoons on the generation prompt — constraining it to the context, forbidding inference, requiring a citation per claim — moved generation from 0.75 to 0.88 and end-to-end to 0.81. The multiplicative model told them that in advance; nobody had written it down.
Evaluating on the set you tuned on. If you adjust chunk size, K, thresholds and prompts against the same 200 questions, you have fitted to those 200 questions. Hold out 30% from the beginning, never look at it during development, and run it once before release. Expect a drop of three to eight points; if it is larger, you overfitted badly.
What this means when you build one
Build the labelled set before you build the pipeline. A hundred questions with known source document IDs, written by someone who actually answers these questions for a living, is the artefact everything else depends on. Written afterwards, it will unconsciously encode the phrasings your system already handles, and it will flatter you.
Then wire the four levels into a single script that prints a decomposed report — recall@50, recall@5, faithfulness, answer relevance, end-to-end correctness, refusal rate on the unanswerable slice — and run it on every change. When a number moves you want to know within a minute which one moved.
Do not ship on an aggregate improvement. Run McNemar's test on the paired results and read the regressions individually; twenty-four questions that got worse is more actionable information than a seven-point aggregate gain, and often more important. Hold 30% of the set back from the start and touch it exactly once, before release.
In production, the two signals worth alerting on from day one are refusal rate and rephrase rate. Both are free, neither needs labels, and between them they catch the two failure modes that actually reach users: the system that has stopped finding things, and the system that has stopped admitting it.
And keep the multiplicative model in your head as the tie-breaker for every prioritisation argument. Whenever someone proposes work on retrieval, ask what the current split is. If generation is the smaller factor, the retrieval work is worth a fraction of what it costs, and the arithmetic will say so before anyone spends a sprint finding out.