Course Content
LLM Evaluation
6 sections · 50 lessons
How do you evaluate a full RAG pipeline end-to-end?
What you need to know
A RAG answer can fail in two different places: the retriever did not find the right passage, or the generator misused a good passage. The fixes are completely different, so the evaluation must separate them.
Retrieval metrics
Context recall — of the facts needed to answer, how many are in the retrieved chunks? RAGAS computes it by splitting the reference answer into claims and checking which are supported by the retrieved context. If the reference has 4 claims and 3 are in the context, context recall = 3 / 4 = 0.75. Recall is a hard ceiling: the generator cannot use evidence it never received.
Context precision — are the relevant chunks ranked near the top? It is rank-aware: average the precision at each position where a relevant chunk appears.
retrieved (top 4): R - R - (R = relevant, - = not relevant)precision@1 = 1/1, precision@3 = 2/3context precision = (1 + 0.667) / 2 = 0.83retrieved (top 4): - - R Rprecision@3 = 1/3, precision@4 = 2/4context precision = (0.333 + 0.5) / 2 = 0.42Both lists contain the same two relevant chunks; the second buries them under noise, which increases cost and the chance the model uses the wrong chunk.
Classic search metrics also work when you have labelled relevant chunks: hit rate@k (was any relevant chunk in the top k?), MRR (mean of 1 / rank of the first relevant chunk: first relevant at rank 3 gives 1/3), and nDCG for graded relevance.
Generation metrics
- Faithfulness — share of the answer's claims supported by the retrieved context. Five claims, four supported: 0.8.
- Answer relevance — does the answer address the question? RAGAS's version generates questions from the answer and measures how similar they are to the original question.
Both are reference-free, so they run on sampled production traffic.
End-to-end metrics
- Answer correctness against gold answers on the golden set.
- Correct abstention on questions the documents do not cover. A RAG system that never says "not in the documents" is not passing.
- Latency split into retrieval and generation, cost per query, and empty-retrieval rate.
The two-by-two diagnosis
| Answer right | Answer wrong | |
|---|---|---|
| Right context retrieved | Working | Generation or prompt problem |
| Wrong context retrieved | Lucky (answered from memory) | Retrieval problem: chunking, embeddings, query, filters |
To confirm, ablate: give the generator the gold context by hand. If answers become right, fix retrieval; if they stay wrong, fix generation.
Tools
RAGAS provides faithfulness, answer relevance, context precision and context recall; DeepEval has equivalent metrics (FaithfulnessMetric, ContextualPrecisionMetric, ContextualRecallMetric); LangSmith and Langfuse run such evaluators on datasets and on production traces. All of them use LLM judges underneath, so calibrate them against human labels on your data.
A real-life example
The HR assistant's golden set has 200 questions, each labelled with the policy sections that answer it, plus 30 questions the policies do not cover. The first full evaluation:
| Metric | Score |
|---|---|
| Context recall (top 5) | 0.71 |
| Context precision | 0.64 |
| Faithfulness | 0.94 |
| Answer correctness | 0.68 |
| Correct abstention (30 uncovered) | 0.53 |
Faithfulness is high, so the generator mostly sticks to its context — the problem is upstream. Feeding gold context raises correctness from 0.68 to 0.89: most errors are retrieval. Error analysis finds two causes: tables in the leave policy were split across chunks, and questions using employee words ("comp off") did not match policy words ("compensatory leave").
The team chunks by policy section, keeps tables whole, and adds a synonym-aware hybrid search (keyword plus embeddings). Context recall rises to 0.88, answer correctness to 0.82. A prompt change — "If the context does not answer the question, say so" — lifts correct abstention to 0.83. These metrics become the CI gate: context recall and correctness may not drop more than 3 points on any pull request that changes chunking, embeddings or prompts.
Follow-up questions to expect
- "Which do you evaluate first, retrieval or generation?" — Retrieval. It is cheaper to measure, it caps everything downstream, and in practice it is where most RAG failures start.
- "How do you get labelled relevant chunks?" — Have domain experts mark the source sections for a few hundred real questions; an LLM can propose candidates for them to confirm. Label sections, not chunk IDs, so labels survive re-chunking.
- "Can faithfulness be high while answers are wrong?" — Yes, when the retriever returns an outdated or wrong document; the answer is faithful to the wrong source. That is why you need correctness on a golden set too.
- "What changes for agentic RAG?" — The system may search several times; evaluate the query-writing steps and the final context set, and add trajectory metrics such as number of searches.