Course Content
LangChain Mastery
7 sections · 109 lessons
How do you evaluate retrieval performance in LangChain?
What you need to know
Two layers of evaluation
Retrieval metrics
- Did the right chunk come back?
- Deterministic, cheap, repeatable
- Needs labelled gold chunk ids
Answer metrics
- Is the final answer correct and supported?
- Needs an LLM judge or human review
- Catches prompt and model problems
If the answer is wrong but the right chunk was retrieved, it is a generation problem. If the right chunk was not retrieved, no prompt can fix it.
The metrics
- recall@k — the share of questions where at least one gold chunk is in the top k. The single most useful number.
- precision@k — how many of the k chunks are relevant; low precision means you pay for noise.
- MRR (mean reciprocal rank) — average of 1 / rank of the first correct chunk. Rank 1 gives 1.0, rank 4 gives 0.25.
- nDCG — like MRR but credits several relevant chunks with graded relevance.
A simple recall function
1def recall_at_k(retriever, examples, k=5):2 hits = 03 for ex in examples: # {"q": str, "gold_ids": set}4 got = [d.metadata["id"] for d in retriever.invoke(ex["q"])[:k]]5 hits += any(i in ex["gold_ids"] for i in got)6 return hits / len(examples)Running it in LangSmith
1from langsmith import Client23def target(inputs: dict) -> dict:4 return {"ids": [d.metadata["id"] for d in retriever.invoke(inputs["question"])]}56def recall_at_5(outputs: dict, reference_outputs: dict) -> bool:7 return any(i in reference_outputs["gold_ids"] for i in outputs["ids"][:5])89Client().evaluate(target, data="hr-retrieval-v1", evaluators=[recall_at_5],10 experiment_prefix="chunk800-k5")Each run becomes an experiment you can compare side by side in the UI.
Where the labelled set comes from
- Real user questions from logs (most valuable), with a person marking the correct chunks.
- Synthetic questions: ask an LLM to write a question each chunk answers, then have a human check a sample.
- Keep some "not in the docs" questions to test that the system says "I don't know".
Answer-level metrics
Faithfulness (every claim supported by the retrieved context), answer relevance, and correctness against a reference answer — the framing popularised by the RAGAS library. Judges drift and can be wrong, so keep the deterministic retrieval metrics as your regression gate in CI.
A real-life example
An HR bot's team wants to switch the embedding model. They have 180 labelled questions from real employee chats.
| Setup | recall@5 | MRR |
|---|---|---|
| Current model, 1,000-char chunks | 0.78 | 0.61 |
| New model, same chunks | 0.80 | 0.63 |
| Current model, 600-char chunks + rerank | 0.89 | 0.77 |
The expensive model change gave little; chunking and reranking gave a lot. They ship the chunking change, and add the retrieval experiment to CI so any change that drops recall@5 below 0.85 fails the build.
Follow-up questions to expect
- "How many examples are enough?" — 50 catches big regressions; 200 or more lets you see differences of a few points.
- "What if you have no gold labels?" — Start with synthetic questions and LLM relevance judgements, then replace them with human labels on real traffic over time.
- "How do you evaluate in production?" — Sample live traces, run faithfulness judges on them, and track user feedback such as thumbs-down rates.