Course Content
LLM Evaluation
6 sections · 50 lessons
How do you evaluate multi-component AI systems holistically?
What you need to know
Real AI products are pipelines: clean the input, classify it, retrieve something, generate, check, act. Each stage can pass its own tests while the product fails.
Level 1: components
Each stage gets its own labelled data and metric:
| Stage | Metric |
|---|---|
| Language detection and clean-up | Accuracy on labelled messages |
| Intent or complaint classifier | Per-class precision, recall, F1 |
| Retriever | Context recall at k, context precision |
| Field extractor | Exact match per field |
| Reply generator | Faithfulness, rubric checks |
These are fast, cheap and precise, and they tell you which stage changed.
Level 2: interfaces
Component tests usually use clean inputs. In production, each stage receives the previous stage's imperfect output. Errors compound:
end-to-end success ≈ product of stage success rates (if errors are independent)0.95 × 0.95 × 0.95 × 0.95 ≈ 0.81Four stages that each look excellent give a system that fails one request in five. Interface tests feed each stage realistic upstream output: a misclassified complaint, a retrieval result with one wrong chunk, a tool that returns an error.
Level 3: the system
Measure the outcome: task success, correct final action, latency and cost per completed task, and safety behaviour of the whole pipeline — on realistic end-to-end scenarios, including rare and adversarial ones. Tie every stage together with one trace ID, so a failed end-to-end case can be walked back to the stage that caused it.
Finding the bottleneck with oracle ablation
- Baseline — run the full system on the end-to-end golden set.
- Swap in one oracle — for example, give the generator the gold retrieval.
- Measure the gain — end-to-end score with the oracle minus the baseline.
- Repeat per component — and invest where the gain is largest.
Keeping it running
Each level has a place in the workflow: component suites in CI on every change, interface and end-to-end suites nightly and before release, and online system metrics (task success, escalations, cost per success) in production, with A/B tests for user-facing changes.
A real-life example
The bank's complaint pipeline has four stages: clean-up (language detection and removal of signatures), classification, policy retrieval, and reply drafting. Component scores all look healthy: clean-up 97%, classification 88%, retrieval recall 85%, drafting faithfulness 95%. Yet only 64% of complaints end with a reply that a supervisor approves unchanged.
The team runs oracle ablations on 300 end-to-end cases:
| Oracle swapped in | End-to-end approval | Gain |
|---|---|---|
| None (baseline) | 64% | — |
| Perfect clean-up | 66% | +2 |
| Perfect classification | 70% | +6 |
| Perfect retrieval | 79% | +15 |
| Perfect classification and retrieval | 86% | +22 |
Retrieval is the biggest lever — and part of its problem is interface-level: when the classifier picks the wrong category, retrieval searches the wrong policy set. The team makes retrieval search the top two predicted categories instead of one, and improves chunking of the card-dispute policy. End-to-end approval reaches 75%. Classification is next on the list. Each stage keeps its own CI suite, and the 300-case end-to-end set runs nightly as the release gate.
Follow-up questions to expect
- "Why not only measure end to end?" — It tells you something broke, not what; with component metrics and traces, a drop can be traced to one stage in minutes.
- "Are stage errors really independent?" — Often not: a hard input tends to fail several stages together, and one stage's error can cause the next. That is why interface tests and oracle ablations beat multiplying component scores.
- "How do you get gold outputs for every stage?" — Label the end-to-end golden set at each stage (category, relevant sections, required reply facts); it is extra work once, and it powers every ablation afterwards.
- "What about cost and latency?" — Measure them per stage and per completed task; a stage that adds 2 seconds for a 1-point gain may not be worth it.