Course Content
RAG Systems
12 sections · 66 lessons
What metrics are used to measure retrieval quality?
What you need to know
To compute any of these, each test question needs a list of relevant chunk ids, ideally with grades (2 = answers it, 1 = helps, 0 = not relevant).
A worked example
A bank's FAQ bot gets "What is the foreclosure charge on a home loan?" The labelled relevant chunks are c2 (grade 2, states the charge), and c4 and c8 (grade 1, related conditions). The retriever returns five chunks:
rank: 1 2 3 4 5chunk: c7 c2 c9 c4 c1grade: 0 2 0 1 0- Hit rate@5 = 1, because at least one relevant chunk is in the top 5.
- Recall@5 = relevant found ÷ total relevant = 2 ÷ 3 = 0.67.
c8was missed. - Precision@5 = relevant found ÷ k = 2 ÷ 5 = 0.4. Three of five chunks are noise.
- MRR (mean reciprocal rank) = 1 ÷ rank of the first relevant chunk = 1 ÷ 2 = 0.5, averaged over all questions.
- nDCG@5 compares the actual order with the perfect order, giving less credit lower down. Here it is 0.54.
The code
1import math23def recall_at_k(ranked, relevant, k):4 return len(set(ranked[:k]) & set(relevant)) / len(relevant)56def mrr(ranked, relevant):7 for rank, d in enumerate(ranked, start=1):8 if d in relevant:9 return 1 / rank10 return 0.01112def ndcg_at_k(ranked, grades, k):13 dcg = sum(grades.get(d, 0) / math.log2(i + 2) for i, d in enumerate(ranked[:k]))14 ideal = sorted(grades.values(), reverse=True)[:k]15 return dcg / sum(g / math.log2(i + 2) for i, g in enumerate(ideal))1617ranked = ["c7", "c2", "c9", "c4", "c1"]18grades = {"c2": 2, "c4": 1, "c8": 1}19print(recall_at_k(ranked, grades, 5), mrr(ranked, grades), ndcg_at_k(ranked, grades, 5))20# 0.666..., 0.5, 0.540...The nDCG function divides each grade by log2(position + 1), so a chunk at rank 1 counts fully and one at rank 4 counts about 43%. Dividing by the score of the ideal order puts the result between 0 and 1.
Which metric for which question
| Metric | Use it when | Position-aware |
|---|---|---|
| Hit rate@k | One chunk is enough to answer | No |
| Recall@k | Answers need several chunks; the ceiling check | No |
| Precision@k | You care about noise and prompt size | No |
| MRR | One right answer, and it should be near the top | Yes |
| nDCG@k | Several chunks with different usefulness | Yes |
Hit rate and recall are often confused. They are the same only when each question has exactly one relevant chunk.
Metrics without chunk labels
RAGAS-style metrics use an LLM judge instead of labelled ids. Context recall checks what share of the reference answer's statements can be found in the retrieved context. Context precision checks whether the useful chunks are ranked near the top. They are useful when labelling chunk ids is too costly, but they inherit the judge's errors.
A real-life example
The bank's team computes these metrics on 300 labelled questions for two retrievers:
| Retriever | Hit rate@5 | Recall@20 | MRR |
|---|---|---|---|
| Dense only | 0.78 | 0.88 | 0.55 |
| Hybrid + reranker | 0.93 | 0.95 | 0.79 |
Reading the pair: dense-only recall@20 of 0.88 says the right chunk is often somewhere in the top 20, but MRR of 0.55 says it is often not first. That points to ranking, not missing content, so a reranker is the right fix. The remaining misses are product codes like HL-FLX-02, which the added BM25 keyword search now catches.
Follow-up questions to expect
- "Which single metric would you watch?" — Recall@k at the k you send to the model, because it caps the whole system.
- "How do you get relevance labels cheaply?" — Start from logged questions and the chunks users' accepted answers cited, have domain experts confirm them, and use an LLM to propose labels that humans review.
- "Why is recall high but answers still wrong?" — Retrieval is fine; look at generation (faithfulness), or at precision, because noise can distract the model.