Course Content
LLM Evaluation
6 sections · 50 lessons
Why should AI evaluation be treated as a system rather than just a model?
What you need to know
The parts of an eval system
dataset (version, hash) -> system under test (prompt, model, params) -> raw outputs (logged) -> graders (rules, judge model + prompt + version) -> aggregation (averages, CIs, thresholds) -> report / CI gateA change anywhere on that line moves the final number.
What "treat it as a system" means
- Version everything — dataset hash, rubric version, judge model ID, judge prompt hash, harness code SHA. Stamp them on every result row.
- Log raw outputs, not just scores, so you can re-grade with a new judge without re-running the system.
- Evaluate the evaluator — a human-labelled calibration set, re-checked on a schedule and whenever the judge changes.
- Monitor the pipeline — judge parse-failure rate, missing rows, score distribution, judge latency and cost. A spike in "could not parse" silently changes averages.
- Test the harness — unit tests for scoring code; a known-good and known-bad output that must always pass and fail.
- Own it — code review for eval changes, and someone on call when it breaks.
A quick diagnosis rule
When a score changes, first ask: did the system change, the data change, or the grader change? If you cannot answer from the stamped versions, you cannot trust the change.
A real-life example
The HR assistant's faithfulness score fell from 0.91 to 0.84 overnight. The team spent two days re-reading prompts and retrieval logs. Nothing had changed. Then someone noticed the judge used a floating model alias, and the provider had moved it to a newer model that was stricter about paraphrased dates.
The fix was structural: the judge is now pinned to a dated model version, every result row records the judge version and prompt hash, and a 150-item calibration set runs whenever anyone proposes a judge change. When they later upgraded the judge on purpose, they re-scored the last month's logged outputs with both judges, published the offset, and started a new baseline — no false alarm.
Follow-up questions to expect
- "How do you validate a new judge version?" — Run old and new judges on the calibration set and on recent logged outputs; compare agreement with humans and the score shift, then reset baselines if needed.
- "What would you put on an eval-health dashboard?" — Judge parse failures, human-agreement over time, score distributions, row counts, judge cost and latency.
- "Who owns the eval?" — The team that owns the product, with the same review standards as production code; evals are part of the product, not a side script.