RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What metrics are used to measure retrieval quality?


Five chunks returned for the foreclosure-charge questionc7· 0c2· 2c9· 0c4· 1c1· 001234firstrelevant: MRR = 1/2partly relevantc8 (grade 1) never appears: recall@5 = 2/3, precision@5 = 2/5, nDCG@5 = 0.54.
Each metric reads the same ranked list differently — recall asks what was missed, MRR asks how high the first hit sits.

What you need to know

To compute any of these, each test question needs a list of relevant chunk ids, ideally with grades (2 = answers it, 1 = helps, 0 = not relevant).

A worked example

A bank's FAQ bot gets "What is the foreclosure charge on a home loan?" The labelled relevant chunks are c2 (grade 2, states the charge), and c4 and c8 (grade 1, related conditions). The retriever returns five chunks:

Text
rank:     1    2    3    4    5chunk:   c7   c2   c9   c4   c1grade:    0    2    0    1    0
  • Hit rate@5 = 1, because at least one relevant chunk is in the top 5.
  • Recall@5 = relevant found ÷ total relevant = 2 ÷ 3 = 0.67. c8 was missed.
  • Precision@5 = relevant found ÷ k = 2 ÷ 5 = 0.4. Three of five chunks are noise.
  • MRR (mean reciprocal rank) = 1 ÷ rank of the first relevant chunk = 1 ÷ 2 = 0.5, averaged over all questions.
  • nDCG@5 compares the actual order with the perfect order, giving less credit lower down. Here it is 0.54.

The code

Python
import mathdef recall_at_k(ranked, relevant, k):    return len(set(ranked[:k]) & set(relevant)) / len(relevant)def mrr(ranked, relevant):    for rank, d in enumerate(ranked, start=1):        if d in relevant:            return 1 / rank    return 0.0def ndcg_at_k(ranked, grades, k):    dcg = sum(grades.get(d, 0) / math.log2(i + 2) for i, d in enumerate(ranked[:k]))    ideal = sorted(grades.values(), reverse=True)[:k]    return dcg / sum(g / math.log2(i + 2) for i, g in enumerate(ideal))ranked = ["c7", "c2", "c9", "c4", "c1"]grades = {"c2": 2, "c4": 1, "c8": 1}print(recall_at_k(ranked, grades, 5), mrr(ranked, grades), ndcg_at_k(ranked, grades, 5))# 0.666..., 0.5, 0.540...

The nDCG function divides each grade by log2(position + 1), so a chunk at rank 1 counts fully and one at rank 4 counts about 43%. Dividing by the score of the ideal order puts the result between 0 and 1.

Which metric for which question

MetricUse it whenPosition-aware
Hit rate@kOne chunk is enough to answerNo
Recall@kAnswers need several chunks; the ceiling checkNo
Precision@kYou care about noise and prompt sizeNo
MRROne right answer, and it should be near the topYes
nDCG@kSeveral chunks with different usefulnessYes

Hit rate and recall are often confused. They are the same only when each question has exactly one relevant chunk.

Metrics without chunk labels

RAGAS-style metrics use an LLM judge instead of labelled ids. Context recall checks what share of the reference answer's statements can be found in the retrieved context. Context precision checks whether the useful chunks are ranked near the top. They are useful when labelling chunk ids is too costly, but they inherit the judge's errors.

A real-life example

The bank's team computes these metrics on 300 labelled questions for two retrievers:

RetrieverHit rate@5Recall@20MRR
Dense only0.780.880.55
Hybrid + reranker0.930.950.79

Reading the pair: dense-only recall@20 of 0.88 says the right chunk is often somewhere in the top 20, but MRR of 0.55 says it is often not first. That points to ranking, not missing content, so a reranker is the right fix. The remaining misses are product codes like HL-FLX-02, which the added BM25 keyword search now catches.

Follow-up questions to expect

  • "Which single metric would you watch?" — Recall@k at the k you send to the model, because it caps the whole system.
  • "How do you get relevance labels cheaply?" — Start from logged questions and the chunks users' accepted answers cited, have domain experts confirm them, and use an LLM to propose labels that humans review.
  • "Why is recall high but answers still wrong?" — Retrieval is fine; look at generation (faithfulness), or at precision, because noise can distract the model.