LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

Why should AI evaluation be treated as a system rather than just a model?


What you need to know

The parts of an eval system

Text
dataset (version, hash)  ->  system under test (prompt, model, params)      -> raw outputs (logged)  ->  graders (rules, judge model + prompt + version)      -> aggregation (averages, CIs, thresholds)  ->  report / CI gate

A change anywhere on that line moves the final number.

What "treat it as a system" means

  • Version everything — dataset hash, rubric version, judge model ID, judge prompt hash, harness code SHA. Stamp them on every result row.
  • Log raw outputs, not just scores, so you can re-grade with a new judge without re-running the system.
  • Evaluate the evaluator — a human-labelled calibration set, re-checked on a schedule and whenever the judge changes.
  • Monitor the pipeline — judge parse-failure rate, missing rows, score distribution, judge latency and cost. A spike in "could not parse" silently changes averages.
  • Test the harness — unit tests for scoring code; a known-good and known-bad output that must always pass and fail.
  • Own it — code review for eval changes, and someone on call when it breaks.

A quick diagnosis rule

When a score changes, first ask: did the system change, the data change, or the grader change? If you cannot answer from the stamped versions, you cannot trust the change.

A real-life example

The HR assistant's faithfulness score fell from 0.91 to 0.84 overnight. The team spent two days re-reading prompts and retrieval logs. Nothing had changed. Then someone noticed the judge used a floating model alias, and the provider had moved it to a newer model that was stricter about paraphrased dates.

The fix was structural: the judge is now pinned to a dated model version, every result row records the judge version and prompt hash, and a 150-item calibration set runs whenever anyone proposes a judge change. When they later upgraded the judge on purpose, they re-scored the last month's logged outputs with both judges, published the offset, and started a new baseline — no false alarm.

Follow-up questions to expect

  • "How do you validate a new judge version?" — Run old and new judges on the calibration set and on recent logged outputs; compare agreement with humans and the score shift, then reset baselines if needed.
  • "What would you put on an eval-health dashboard?" — Judge parse failures, human-agreement over time, score distributions, row counts, judge cost and latency.
  • "Who owns the eval?" — The team that owns the product, with the same review standards as production code; evals are part of the product, not a side script.