LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you evaluate retrieval performance in LangChain?


One labelled set, three candidate setupsCurrent, 1,000 chars0.780.61New embedding model0.800.63600 charsplus rerank0.890.77Setuprecall at 5MRR
The expensive model swap moved recall two points; chunking and reranking moved it eleven — only a fixed set shows that.

What you need to know

Two layers of evaluation

Retrieval metrics

  • Did the right chunk come back?
  • Deterministic, cheap, repeatable
  • Needs labelled gold chunk ids

Answer metrics

  • Is the final answer correct and supported?
  • Needs an LLM judge or human review
  • Catches prompt and model problems

If the answer is wrong but the right chunk was retrieved, it is a generation problem. If the right chunk was not retrieved, no prompt can fix it.

The metrics

  • recall@k — the share of questions where at least one gold chunk is in the top k. The single most useful number.
  • precision@k — how many of the k chunks are relevant; low precision means you pay for noise.
  • MRR (mean reciprocal rank) — average of 1 / rank of the first correct chunk. Rank 1 gives 1.0, rank 4 gives 0.25.
  • nDCG — like MRR but credits several relevant chunks with graded relevance.

A simple recall function

Python
def recall_at_k(retriever, examples, k=5):    hits = 0    for ex in examples:                              # {"q": str, "gold_ids": set}        got = [d.metadata["id"] for d in retriever.invoke(ex["q"])[:k]]        hits += any(i in ex["gold_ids"] for i in got)    return hits / len(examples)

Running it in LangSmith

Python
from langsmith import Clientdef target(inputs: dict) -> dict:    return {"ids": [d.metadata["id"] for d in retriever.invoke(inputs["question"])]}def recall_at_5(outputs: dict, reference_outputs: dict) -> bool:    return any(i in reference_outputs["gold_ids"] for i in outputs["ids"][:5])Client().evaluate(target, data="hr-retrieval-v1", evaluators=[recall_at_5],                  experiment_prefix="chunk800-k5")

Each run becomes an experiment you can compare side by side in the UI.

Where the labelled set comes from

  • Real user questions from logs (most valuable), with a person marking the correct chunks.
  • Synthetic questions: ask an LLM to write a question each chunk answers, then have a human check a sample.
  • Keep some "not in the docs" questions to test that the system says "I don't know".

Answer-level metrics

Faithfulness (every claim supported by the retrieved context), answer relevance, and correctness against a reference answer — the framing popularised by the RAGAS library. Judges drift and can be wrong, so keep the deterministic retrieval metrics as your regression gate in CI.

A real-life example

An HR bot's team wants to switch the embedding model. They have 180 labelled questions from real employee chats.

Setuprecall@5MRR
Current model, 1,000-char chunks0.780.61
New model, same chunks0.800.63
Current model, 600-char chunks + rerank0.890.77

The expensive model change gave little; chunking and reranking gave a lot. They ship the chunking change, and add the retrieval experiment to CI so any change that drops recall@5 below 0.85 fails the build.

Follow-up questions to expect

  • "How many examples are enough?" — 50 catches big regressions; 200 or more lets you see differences of a few points.
  • "What if you have no gold labels?" — Start with synthetic questions and LLM relevance judgements, then replace them with human labels on real traffic over time.
  • "How do you evaluate in production?" — Sample live traces, run faithfulness judges on them, and track user feedback such as thumbs-down rates.