Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Evaluating RAG Performance


An engineering team stood up in a review and said their RAG system was 92% accurate. Someone asked what that meant. It turned out to mean this: on a set of 200 test questions, the correct source document appeared in the top 5 retrieved chunks 184 times.

Support had a different number. They had sampled 60 real conversations that week and judged 21 of the answers wrong or misleading — about 65% correct. Both numbers were honestly obtained. Both were correct. They were measuring different things, and neither told anybody what to fix.

When they finally instrumented it properly, the picture was unambiguous. Of the 184 questions where the right document was retrieved, the generator produced a correct answer only 138 times. End to end: 138 out of 200, or 69%. The retriever was fine. The generator was throwing away one answer in four that had been handed to it on a plate.

That is the entire case for structured evaluation. A single number cannot tell you which half of a two-stage system is broken, and RAG has at least two halves. You need numbers that decompose.

Four levels, evaluated separatelyRetrieval: recall at k, MRR, NDCG at kGeneration: faithfulness to the given contextEnd to end: does it answer the question askedProduction: latency, cost, no-answer rate, drift
'92 percent accurate' meant recall at 5 on 200 questions — a level-one number quoted as if it described the whole system.

A framework that decomposes

Evaluate at four levels, each answering a question the level below cannot.

LevelQuestion it answersNeeds labels?Run it
1. RetrievalDid we find the right documents, and rank them well?Yes — relevant doc IDs per questionEvery commit
2. GenerationGiven those documents, was the answer grounded and on-point?Partly — LLM judge plus a human-labelled calibration setEvery commit
3. End-to-endWould a user call this answer correct?Yes — gold answersBefore every release
4. ProductionIs it still working on traffic we never anticipated?No — implicit signals onlyContinuously

The relationship between levels 1, 2 and 3 is roughly multiplicative:

P(correct answer)≈P(retrieved)×P(correct∣retrieved)P(\text{correct answer}) \approx P(\text{retrieved}) \times P(\text{correct} \mid \text{retrieved})

For the team above: 0.92×0.75=0.690.92 \times 0.75 = 0.69. That equation is the most useful thing in this lesson, because it tells you where your headroom is. Pushing retrieval from 0.92 to 0.97 buys 0.97×0.75=0.730.97 \times 0.75 = 0.73 — four points. Pushing generation from 0.75 to 0.90 buys 0.92×0.90=0.830.92 \times 0.90 = 0.83 — fourteen points, for less work. Teams spend their time on retrieval because retrieval is the fun part.

Measure both factors separately, then work on the smaller one. A system with 0.95 retrieval and 0.60 generation is not a retrieval problem, however much it feels like one.

Level 1: retrieval evaluation

You need a labelled set: for each question, the IDs of the documents that genuinely contain the answer. Two hundred questions is enough to start. Building it is a day of tedious work and it is the highest-return day you will spend on the project.

Take one question with three relevant documents in the corpus — d17, d42, d88 — landing at ranks 1, 3 and 12.

Precision@K and Recall@K

Precision@K=∣relevant∩retrieved@K∣KRecall@K=∣relevant∩retrieved@K∣∣relevant∣\text{Precision@K} = \frac{|\text{relevant} \cap \text{retrieved@K}|}{K} \qquad \text{Recall@K} = \frac{|\text{relevant} \cap \text{retrieved@K}|}{|\text{relevant}|}

KRelevant foundPrecision@KRecall@K
11 (d17)1/1 = 1.0001/3 = 0.333
32 (d17, d42)2/3 = 0.6672/3 = 0.667
522/5 = 0.4002/3 = 0.667
1022/10 = 0.2002/3 = 0.667
203 (all)3/20 = 0.1503/3 = 1.000

Read the two columns against each other. Precision falls monotonically as K grows; recall rises. There is no K that is good at both, so you have to decide which failure you can live with.

For RAG, recall matters more, and for a specific reason: a document you did not retrieve is unrecoverable — nothing downstream can invent it — whereas an irrelevant document is merely noise a re-ranker or a well-constrained generator can survive.

But precision is not free either, and this is where people over-correct. Push K to 20 and you are sending 20 chunks to the generator. Models measurably degrade when the relevant fact is buried in the middle of a long context, and irrelevant chunks raise hallucination rates. The standard resolution is to retrieve wide and re-rank narrow: high recall at K=50, high precision at K=5.

MRR — when there is one right answer

MRR=1∣Q∣∑i=1∣Q∣1ranki\text{MRR} = \frac{1}{|Q|}\sum_{i=1}^{|Q|}\frac{1}{\text{rank}_i}

Take the reciprocal of the rank of the first relevant document. Across four questions with first-relevant ranks 1, 4, 2 and "not found in top 10", the reciprocal ranks are 1.000, 0.250, 0.500 and 0.000, giving MRR=1.75/4=0.4375\text{MRR} = 1.75/4 = 0.4375.

MRR is the right metric for factoid lookup, where one document answers the question and you only care how near the top it lands. It is the wrong metric when several documents each contribute part of the answer, because it ignores everything after the first hit.

NDCG@K — when relevance comes in grades

NDCG handles graded relevance (3 = directly answers, 2 = useful context, 1 = tangential, 0 = irrelevant) and discounts each hit by its position.

DCG@K=∑i=1Krelilog⁡2(i+1)NDCG@K=DCG@KIDCG@K\text{DCG@K} = \sum_{i=1}^{K}\frac{rel_i}{\log_2(i+1)} \qquad \text{NDCG@K} = \frac{\text{DCG@K}}{\text{IDCG@K}}

Grades [2, 0, 3, 1, 0] at ranks 1-5 give

DCG@5=21.0000+0+32.0000+12.3219+0=2.0000+1.5000+0.4307=3.9307\text{DCG@5} = \frac{2}{1.0000} + 0 + \frac{3}{2.0000} + \frac{1}{2.3219} + 0 = 2.0000 + 1.5000 + 0.4307 = 3.9307

The ideal ordering of those same grades is [3, 2, 1, 0, 0], so

IDCG@5=3.0000+21.5850+12.0000=3.0000+1.2619+0.5000=4.7619\text{IDCG@5} = 3.0000 + \frac{2}{1.5850} + \frac{1}{2.0000} = 3.0000 + 1.2619 + 0.5000 = 4.7619

and NDCG@5=3.9307/4.7619=0.825\text{NDCG@5} = 3.9307 / 4.7619 = 0.825. Two warnings: an exponential-gain variant using 2reli−12^{rel_i}-1 is equally common and gives different numbers, so always state which you used; and grading relevance on a four-point scale needs an annotation guideline, or two people will disagree on half the labels.

MAP@K — precision averaged over every hit

Average Precision takes Precision@i at each rank ii where a relevant document appears, and averages. For our three documents at ranks 1, 3 and 12:

Relevant doc at rankPrecision at that rankValue
11/11.000
32/30.667
123/120.250

AP=(1.000+0.667+0.250)/3=1.917/3=0.639\text{AP} = (1.000 + 0.667 + 0.250)/3 = 1.917/3 = 0.639. For a second question with relevant documents at ranks 2 and 5, AP=(1/2+2/5)/2=0.900/2=0.450\text{AP} = (1/2 + 2/5)/2 = 0.900/2 = 0.450. Across the two, MAP=(0.639+0.450)/2=0.544\text{MAP} = (0.639 + 0.450)/2 = \mathbf{0.544}.

MAP is the metric to report when multiple documents matter and you want one number that respects both how many you found and where you put them.

UseWhenBlind to
Recall@KChoosing your fetch depth before re-rankingOrder entirely
Precision@KSetting the final K sent to the generatorWhat you missed
MRRSingle-answer factoid lookupEverything after the first hit
NDCG@KGraded relevance; comparing re-rankersNothing much — but needs graded labels
MAP@KMultiple relevant docs, binary labelsDegrees of relevance

Level 2: generation evaluation

Now hold retrieval fixed. Feed the generator the correct context and ask what it did with it.

Faithfulness — is every claim supported by the context?

This is the metric that catches hallucination, and the method is to decompose the answer into atomic claims and check each one against the retrieved text.

Take a real answer:

"Enterprise annual contracts have a 14-day refund window from the invoice date. Refunds require approval from two signatories. Partial refunds are prorated monthly. The policy was last updated in March 2024."

Split it and check:

#Atomic claimIn context?
1Enterprise annual contracts have a 14-day refund windowYes
2The window runs from the invoice dateYes
3Refunds require two signatoriesYes
4Partial refunds are prorated monthlyNo
5The policy was last updated in March 2024No

Faithfulness =3/5=0.60= 3/5 = \mathbf{0.60}. Look closely at claims 4 and 5. They are not wild inventions — they are exactly the kind of plausible administrative detail a model adds because policy documents usually contain such things. That is what real hallucination looks like in production: not nonsense, but confident specifics in the same register as the true parts, which is precisely why humans reading the answer do not notice.

Python
FAITHFULNESS = """You are checking whether a claim is supported by context.Context:{context}Claim: {claim}Answer with exactly one word: SUPPORTED, CONTRADICTED, or NOT_STATED.A claim is SUPPORTED only if the context states it or entails it directly.Plausible inference is NOT support."""def faithfulness(answer, context, judge, splitter):    claims = splitter(answer)              # one factual assertion each    verdicts = [judge.invoke(FAITHFULNESS.format(        context=context, claim=c)).content.strip() for c in claims]    return sum(v == "SUPPORTED" for v in verdicts) / len(claims)

The last line of that prompt is doing real work. Without it, judges mark anything reasonable as supported, and faithfulness scores cluster around 0.95 regardless of how the system is behaving.

Calibrate the judge, or the judge is decoration

An LLM judge is a model, and unvalidated models lie. Label 100 claim-context pairs by hand, run the judge over the same 100, and compute agreement. Raw agreement is misleading when one class dominates, so use Cohen's kappa:

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

Suppose the judge and the human agree on 88 of 100 pairs, so po=0.88p_o = 0.88. The human marked 70 supported, the judge 76, so chance agreement is pe=(0.70)(0.76)+(0.30)(0.24)=0.532+0.072=0.604p_e = (0.70)(0.76) + (0.30)(0.24) = 0.532 + 0.072 = 0.604. Then κ=(0.88−0.604)/(1−0.604)=0.276/0.396=0.697\kappa = (0.88 - 0.604)/(1 - 0.604) = 0.276/0.396 = \mathbf{0.697}.

Kappa around 0.70 is substantial agreement and usable for tracking relative change. Below about 0.60, your judge is not measuring what you think, and every conclusion you draw from it is noise. Re-check kappa whenever you change the judge model or the prompt.

Answer relevance — does it address the question asked?

An answer can be perfectly faithful and completely beside the point. The standard technique runs backwards: ask a model to generate the questions this answer would be a good response to, embed them, and compare each to the real question.

For the question "What is the refund window for enterprise annual contracts?", the generated questions and their cosine similarities might be 0.91, 0.88 and 0.62, giving relevance =(0.91+0.88+0.62)/3=2.41/3=0.803= (0.91+0.88+0.62)/3 = 2.41/3 = \mathbf{0.803}. The 0.62 is the informative one — it usually corresponds to a paragraph where the answer drifted into a related but unasked topic, which is a retrieval precision problem showing up as a generation symptom.

Fluency, conciseness and context utilisation

Modern models are fluent by default, so fluency is rarely your bottleneck and rarely worth measuring. Two cheaper proxies earn their place:

  • Length ratio — answer tokens divided by gold answer tokens. Consistently above about 2.5 means the model is padding, restating the question, and hedging, all of which users read as evasive.
  • Context utilisation — the fraction of retrieved chunks actually cited in the answer. Send 5 chunks, cite 1, and you are paying five times over for one document's worth of value. Utilisation below 0.4 is a signal to reduce K, not to improve the prompt.

Level 3: end-to-end evaluation

Levels 1 and 2 diagnose. Level 3 decides whether to ship. It needs gold answers, and the composition of that set matters more than its size.

SliceShareWhy it must be there
Single-fact lookup40%The bulk of real traffic
Multi-document synthesis20%Catches the retriever returning one facet of a many-part answer
Comparison and conditionals15%Where negation and qualifiers break both stages
Ambiguous or under-specified10%Tests whether the system asks rather than guesses
Unanswerable15%The only way to measure whether it refuses

That last row is the one everybody omits, and omitting it is the reason so many systems score well in evaluation and hallucinate in production. If every question in your test set has an answer in the corpus, a system that never refuses scores perfectly — and you have accidentally optimised for confident guessing. Include questions you know the corpus cannot answer and score the refusal as the correct response.

A/B testing two configurations, honestly

Run both configurations over the same questions and compare per question, not in aggregate. Suppose Config A gets 138 of 200 right and Config B gets 152 — a seven-point improvement. Ship it?

Not yet. Look at the paired outcomes:

B correctB wrong
A correct11424
A wrong3824

The 114 and the 24 in the corner tell you nothing — both configurations behaved identically there. The evidence lives entirely in the discordant cells: B fixed 38 questions A got wrong, and broke 24 that A got right. McNemar's test uses exactly those two numbers:

χ2=(∣b−c∣−1)2b+c=(∣24−38∣−1)224+38=13262=16962=2.726\chi^2 = \frac{(|b - c| - 1)^2}{b + c} = \frac{(|24 - 38| - 1)^2}{24 + 38} = \frac{13^2}{62} = \frac{169}{62} = 2.726

The critical value for one degree of freedom at the 5% level is 3.841. Since 2.726<3.8412.726 < 3.841 (p ≈ 0.10), this result is not statistically significant. A seven-point improvement on 200 questions is well inside what you would expect from shuffling which questions happen to be in your test set.

Twenty-four questions got worse. If you ship on the aggregate number alone you will never look at them, and one of them will be the query your largest customer runs every morning.

If the same effect held at 500 questions — 95 fixed, 60 broken — then χ2=(35−1)2/155=1156/155=7.46\chi^2 = (35-1)^2/155 = 1156/155 = 7.46, comfortably above 3.841. Either collect more questions or accept that you cannot yet tell the two configurations apart. And read the 24 regressions individually: they usually cluster into one recognisable pattern, and that pattern is more useful than the aggregate ever was.

Level 4: production monitoring

Offline sets go stale within weeks. Real users ask things nobody anticipated, in phrasing nobody wrote down. The trick in production is that you have no labels — so you monitor signals that correlate with failure without needing them.

SignalWhat a change meansAct when
Median retrieval top scoreCorpus drifting away from query distributionDrops more than 0.05 week on week
Score margin (1st − 5th)Flat profiles mean the corpus has nothing specificMedian below 0.15
Refusal rateRising: coverage gaps. Falling sharply: guardrail brokeMoves more than 5 points either way
Rephrase rateUser re-asks the same thing within 60 s — the answer failedAbove baseline + 5 points
Citation rateFalling means answers coming from parameters, not contextBelow 0.85
Escalation to humanThe most honest quality signal you haveAny sustained rise
p95 latencyUsers abandon before quality mattersAbove your stated budget

Rephrase rate deserves the emphasis. If a user asks a semantically similar question within a minute of the last one, the first answer almost certainly failed them. It needs no labelling, no judge and no annotation budget — just an embedding comparison between consecutive queries in a session. One team's baseline sat at 8%; a deploy pushed it to 19% overnight, and the alert fired eleven hours before the first support ticket arrived.

Sample and label a hundred real conversations a week regardless. Automated signals tell you that something changed; only reading transcripts tells you what.

Reading the numbers back to a fix

The point of measuring in layers is that each combination of results points at one stage.

Recall@50Recall@5FaithfulnessDiagnosisFix
Low (<0.85)Low—The answer is not in the candidate set at allChunking, embedding model, add BM25 — a re-ranker cannot help
HighLow—Found but badly orderedAdd or upgrade the cross-encoder re-ranker
HighHighLowRight context, ungrounded answerConstrain the prompt to context, temperature to 0, require citations
HighHighHigh, relevance lowRight topic, wrong facetQuery condensation and rewriting; check chunk boundaries
HighHighHighMetrics good, users unhappyFormat, tone, latency, missing citations — measure the right thing

Two failure modes deserve naming because they waste the most time.

Optimising the wrong factor. The team in the opening spent six weeks on embeddings and re-ranking, moving retrieval from 0.89 to 0.92. End-to-end went from 0.67 to 0.69. Two afternoons on the generation prompt — constraining it to the context, forbidding inference, requiring a citation per claim — moved generation from 0.75 to 0.88 and end-to-end to 0.81. The multiplicative model told them that in advance; nobody had written it down.

Evaluating on the set you tuned on. If you adjust chunk size, K, thresholds and prompts against the same 200 questions, you have fitted to those 200 questions. Hold out 30% from the beginning, never look at it during development, and run it once before release. Expect a drop of three to eight points; if it is larger, you overfitted badly.

What this means when you build one

Build the labelled set before you build the pipeline. A hundred questions with known source document IDs, written by someone who actually answers these questions for a living, is the artefact everything else depends on. Written afterwards, it will unconsciously encode the phrasings your system already handles, and it will flatter you.

Then wire the four levels into a single script that prints a decomposed report — recall@50, recall@5, faithfulness, answer relevance, end-to-end correctness, refusal rate on the unanswerable slice — and run it on every change. When a number moves you want to know within a minute which one moved.

Do not ship on an aggregate improvement. Run McNemar's test on the paired results and read the regressions individually; twenty-four questions that got worse is more actionable information than a seven-point aggregate gain, and often more important. Hold 30% of the set back from the start and touch it exactly once, before release.

In production, the two signals worth alerting on from day one are refusal rate and rephrase rate. Both are free, neither needs labels, and between them they catch the two failure modes that actually reach users: the system that has stopped finding things, and the system that has stopped admitting it.

And keep the multiplicative model in your head as the tie-breaker for every prioritisation argument. Whenever someone proposes work on retrieval, ask what the current split is. If generation is the smaller factor, the retrieval work is worth a fraction of what it costs, and the arithmetic will say so before anyone spends a sprint finding out.