LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you ensure evaluation results are reproducible?


What you need to know

Sources of irreproducibility

  • Floating model aliases — a name like "latest" can point to a new model tomorrow. The most common cause of "nothing changed but the score did".
  • Sampling — temperature above 0 gives different outputs by design.
  • Server-side non-determinism — even at temperature 0, batching, floating-point order and mixture-of-experts routing can change outputs. Seeds, where an API offers them, are best effort. Some reasoning models do not accept a temperature at all.
  • Dataset edits — someone fixes a label; historical comparisons silently break.
  • Judge changes — a new judge model or rubric wording.
  • Harness code — a changed answer-extraction regex can move a benchmark score by points.

What to record for every run

JSON
{"run_id": "2026-09-14-nightly", "git_sha": "a91f2c0", "dataset": {"name": "hr-golden", "hash": "sha256:8c1e...", "rows": 200}, "system": {"model": "provider-model-2026-06-01", "prompt_hash": "77ab", "temperature": 0}, "judge": {"model": "judge-model-2026-05-15", "prompt_hash": "4de9"}, "n_repeats": 3}

Every result row carries this, so any two runs can be diffed field by field.

Practices

  • Store raw outputs; re-grading is cheap, re-generating is not.
  • Keep a small canary set that should always give the same verdicts; if they change, something in the pipeline changed.
  • Report mean and confidence interval across repeats, not the best of several runs.
  • Pin library versions and use a container image for the harness.

A real-life example

The code-review bot's nightly eval showed bug-catch recall of 74%, then 69%, then 73% on three nights with no code changes. The team suspected the model provider. The run records told a different story: the dataset hash changed on the second night, because a teammate had "cleaned" 12 PRs, removing some of the harder seeded bugs and then restoring them.

After that, datasets became read-only snapshots with hashes, changes went through code review, and each nightly run was repeated three times. Run-to-run spread fell to about ±1.5 points, which also let them tighten the CI regression margin from 5 points to 3.

Follow-up questions to expect

  • "Is temperature 0 enough for reproducibility?" — No. It reduces variation but does not guarantee identical outputs on hosted APIs; repeats and intervals are still needed.
  • "How do you reproduce an old result after the provider retires the model?" — You can't re-generate; that is why raw outputs are stored. You can re-grade them and compare new models against those stored results.
  • "What about open-weight models?" — More control: pin weights, inference engine version and hardware settings; still expect small differences across GPU types and batch sizes.