Course Content
RAG Systems
12 sections · 66 lessons
How do you evaluate the quality of a RAG system?
What you need to know
A RAG answer can fail in two places: the retriever did not find the right text, or the model did not use it well. These have opposite fixes, so you need separate numbers.
Level 1: retrieval
Given a question, did the right chunks come back, and how high? Metrics: recall@k (share of the relevant chunks found in the top k), hit rate@k (was at least one found), MRR (how high the first relevant one is). These are cheap to compute and need no LLM. They are covered in the next lesson.
Level 2: generation
Given the retrieved chunks, is the answer good?
- Faithfulness: is every claim supported by the retrieved context?
- Answer relevance: does it actually address the question?
- Correctness: does it match the reference answer?
These need judgment, so teams use an LLM-as-judge: a second model call with a scoring rubric. RAGAS, DeepEval, TruLens, LangSmith and Langfuse all provide versions of these metrics.
Level 3: end to end
Task success rate, share of valid citations, refusal rate on answerable questions, p95 latency, and cost per query.
Trusting the judge
An LLM judge has known biases:
- Position bias: when comparing two answers, it tends to prefer one position (for example, the first).
- Verbosity bias: it tends to prefer longer answers.
- Self-preference: it may rate text from its own model family higher.
- Leniency: without a strict rubric, it gives most answers a pass.
So before trusting a judge, have humans label 50 to 100 answers, run the judge on the same answers, and measure agreement. Use a narrow yes/no question per criterion, swap the order in pairwise tests, and use a different model family from the one generating answers where you can.
The workflow
- Build the golden set — real user questions, reference answers and source chunk ids, written or checked by domain experts.
- Run it in CI — on every change to prompt, chunking, embedding model, reranker or LLM.
- Gate the release — a drop beyond an agreed margin on any key metric blocks the deploy.
- Monitor production — cheap checks on every request, judge scores on a sample, user feedback.
- Feed failures back — each confirmed bad answer becomes a new golden-set case.
A real-life example
An HR policy assistant has a golden set of 200 questions written with the HR team. The engineer wants to change the chunk size from 1,000 to 500 characters. The CI report compares the two:
| Metric | Chunk 1,000 | Chunk 500 |
|---|---|---|
| Recall@5 | 0.86 | 0.91 |
| Faithfulness | 0.94 | 0.95 |
| Correctness | 0.81 | 0.78 |
| Input tokens per query | 1,700 | 1,100 |
Retrieval improved, but correctness fell. Reading the failures shows why: leave rules with long tables are now split in half, so answers about grade-specific entitlements lose the header row. Without separate stage metrics, the team would only have seen "correctness fell 3 points" and not known where to look. They keep 500 characters for normal text and split tables as whole units.
Follow-up questions to expect
- "Where do golden-set questions come from?" — Real user logs first, then questions experts write for important edge cases, including questions the corpus cannot answer (to test refusals).
- "Can you generate test questions with an LLM?" — Yes, to bootstrap coverage, but they tend to copy the document's wording and be easier than real questions. Mix them with real ones and review them.
- "How big should the set be?" — Big enough that a real change is larger than the noise. With 200 questions, a change of one or two questions is noise; look for consistent shifts and read the failures.