LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you evaluate multi-component AI systems holistically?


End-to-end approval with one oracle swapped in64%66%70%79%86%01234baselineclean-upclassifierretrievalbothEvery component scored 85% or better on its own tests.
Perfect retrieval buys 15 points and perfect clean-up only 2, so the oracle test says where the next month should go.

What you need to know

Real AI products are pipelines: clean the input, classify it, retrieve something, generate, check, act. Each stage can pass its own tests while the product fails.

Level 1: components

Each stage gets its own labelled data and metric:

StageMetric
Language detection and clean-upAccuracy on labelled messages
Intent or complaint classifierPer-class precision, recall, F1
RetrieverContext recall at k, context precision
Field extractorExact match per field
Reply generatorFaithfulness, rubric checks

These are fast, cheap and precise, and they tell you which stage changed.

Level 2: interfaces

Component tests usually use clean inputs. In production, each stage receives the previous stage's imperfect output. Errors compound:

Text
end-to-end success ≈ product of stage success rates (if errors are independent)0.95 × 0.95 × 0.95 × 0.95 ≈ 0.81

Four stages that each look excellent give a system that fails one request in five. Interface tests feed each stage realistic upstream output: a misclassified complaint, a retrieval result with one wrong chunk, a tool that returns an error.

Level 3: the system

Measure the outcome: task success, correct final action, latency and cost per completed task, and safety behaviour of the whole pipeline — on realistic end-to-end scenarios, including rare and adversarial ones. Tie every stage together with one trace ID, so a failed end-to-end case can be walked back to the stage that caused it.

Finding the bottleneck with oracle ablation

  1. Baseline — run the full system on the end-to-end golden set.
  2. Swap in one oracle — for example, give the generator the gold retrieval.
  3. Measure the gain — end-to-end score with the oracle minus the baseline.
  4. Repeat per component — and invest where the gain is largest.

Keeping it running

Each level has a place in the workflow: component suites in CI on every change, interface and end-to-end suites nightly and before release, and online system metrics (task success, escalations, cost per success) in production, with A/B tests for user-facing changes.

A real-life example

The bank's complaint pipeline has four stages: clean-up (language detection and removal of signatures), classification, policy retrieval, and reply drafting. Component scores all look healthy: clean-up 97%, classification 88%, retrieval recall 85%, drafting faithfulness 95%. Yet only 64% of complaints end with a reply that a supervisor approves unchanged.

The team runs oracle ablations on 300 end-to-end cases:

Oracle swapped inEnd-to-end approvalGain
None (baseline)64%—
Perfect clean-up66%+2
Perfect classification70%+6
Perfect retrieval79%+15
Perfect classification and retrieval86%+22

Retrieval is the biggest lever — and part of its problem is interface-level: when the classifier picks the wrong category, retrieval searches the wrong policy set. The team makes retrieval search the top two predicted categories instead of one, and improves chunking of the card-dispute policy. End-to-end approval reaches 75%. Classification is next on the list. Each stage keeps its own CI suite, and the 300-case end-to-end set runs nightly as the release gate.

Follow-up questions to expect

  • "Why not only measure end to end?" — It tells you something broke, not what; with component metrics and traces, a drop can be traced to one stage in minutes.
  • "Are stage errors really independent?" — Often not: a hard input tends to fail several stages together, and one stage's error can cause the next. That is why interface tests and oracle ablations beat multiplying component scores.
  • "How do you get gold outputs for every stage?" — Label the end-to-end golden set at each stage (category, relevant sections, required reply facts); it is extra work once, and it powers every ablation afterwards.
  • "What about cost and latency?" — Measure them per stage and per completed task; a stage that adds 2 seconds for a 1-point gain may not be worth it.