RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you evaluate the quality of a RAG system?


Three levels, each with its own numbersRetrieval — recall@k, MRR; no LLM neededGeneration — faithfulness,relevance; judge checked vs humansEnd to end — correctness, refusals, p95, costProduction — sampled judge,feedback into the golden set
When chunk size 500 raised recall but lowered correctness, only separate levels showed that tables were being cut in half.

What you need to know

A RAG answer can fail in two places: the retriever did not find the right text, or the model did not use it well. These have opposite fixes, so you need separate numbers.

Level 1: retrieval

Given a question, did the right chunks come back, and how high? Metrics: recall@k (share of the relevant chunks found in the top k), hit rate@k (was at least one found), MRR (how high the first relevant one is). These are cheap to compute and need no LLM. They are covered in the next lesson.

Level 2: generation

Given the retrieved chunks, is the answer good?

  • Faithfulness: is every claim supported by the retrieved context?
  • Answer relevance: does it actually address the question?
  • Correctness: does it match the reference answer?

These need judgment, so teams use an LLM-as-judge: a second model call with a scoring rubric. RAGAS, DeepEval, TruLens, LangSmith and Langfuse all provide versions of these metrics.

Level 3: end to end

Task success rate, share of valid citations, refusal rate on answerable questions, p95 latency, and cost per query.

Trusting the judge

An LLM judge has known biases:

  • Position bias: when comparing two answers, it tends to prefer one position (for example, the first).
  • Verbosity bias: it tends to prefer longer answers.
  • Self-preference: it may rate text from its own model family higher.
  • Leniency: without a strict rubric, it gives most answers a pass.

So before trusting a judge, have humans label 50 to 100 answers, run the judge on the same answers, and measure agreement. Use a narrow yes/no question per criterion, swap the order in pairwise tests, and use a different model family from the one generating answers where you can.

The workflow

  1. Build the golden set — real user questions, reference answers and source chunk ids, written or checked by domain experts.
  2. Run it in CI — on every change to prompt, chunking, embedding model, reranker or LLM.
  3. Gate the release — a drop beyond an agreed margin on any key metric blocks the deploy.
  4. Monitor production — cheap checks on every request, judge scores on a sample, user feedback.
  5. Feed failures back — each confirmed bad answer becomes a new golden-set case.

A real-life example

An HR policy assistant has a golden set of 200 questions written with the HR team. The engineer wants to change the chunk size from 1,000 to 500 characters. The CI report compares the two:

MetricChunk 1,000Chunk 500
Recall@50.860.91
Faithfulness0.940.95
Correctness0.810.78
Input tokens per query1,7001,100

Retrieval improved, but correctness fell. Reading the failures shows why: leave rules with long tables are now split in half, so answers about grade-specific entitlements lose the header row. Without separate stage metrics, the team would only have seen "correctness fell 3 points" and not known where to look. They keep 500 characters for normal text and split tables as whole units.

Follow-up questions to expect

  • "Where do golden-set questions come from?" — Real user logs first, then questions experts write for important edge cases, including questions the corpus cannot answer (to test refusals).
  • "Can you generate test questions with an LLM?" — Yes, to bootstrap coverage, but they tend to copy the document's wording and be easier than real questions. Mix them with real ones and review them.
  • "How big should the set be?" — Big enough that a real change is larger than the noise. With 200 questions, a change of one or two questions is noise; look for consistent shifts and read the failures.