Course Content
LLM Evaluation
6 sections · 50 lessons
How do you ensure evaluation results are reproducible?
What you need to know
Sources of irreproducibility
- Floating model aliases — a name like "latest" can point to a new model tomorrow. The most common cause of "nothing changed but the score did".
- Sampling — temperature above 0 gives different outputs by design.
- Server-side non-determinism — even at temperature 0, batching, floating-point order and mixture-of-experts routing can change outputs. Seeds, where an API offers them, are best effort. Some reasoning models do not accept a temperature at all.
- Dataset edits — someone fixes a label; historical comparisons silently break.
- Judge changes — a new judge model or rubric wording.
- Harness code — a changed answer-extraction regex can move a benchmark score by points.
What to record for every run
JSON
1{"run_id": "2026-09-14-nightly",2 "git_sha": "a91f2c0",3 "dataset": {"name": "hr-golden", "hash": "sha256:8c1e...", "rows": 200},4 "system": {"model": "provider-model-2026-06-01", "prompt_hash": "77ab", "temperature": 0},5 "judge": {"model": "judge-model-2026-05-15", "prompt_hash": "4de9"},6 "n_repeats": 3}Every result row carries this, so any two runs can be diffed field by field.
Practices
- Store raw outputs; re-grading is cheap, re-generating is not.
- Keep a small canary set that should always give the same verdicts; if they change, something in the pipeline changed.
- Report mean and confidence interval across repeats, not the best of several runs.
- Pin library versions and use a container image for the harness.
A real-life example
The code-review bot's nightly eval showed bug-catch recall of 74%, then 69%, then 73% on three nights with no code changes. The team suspected the model provider. The run records told a different story: the dataset hash changed on the second night, because a teammate had "cleaned" 12 PRs, removing some of the harder seeded bugs and then restoring them.
After that, datasets became read-only snapshots with hashes, changes went through code review, and each nightly run was repeated three times. Run-to-run spread fell to about ±1.5 points, which also let them tighten the CI regression margin from 5 points to 3.
Follow-up questions to expect
- "Is temperature 0 enough for reproducibility?" — No. It reduces variation but does not guarantee identical outputs on hosted APIs; repeats and intervals are still needed.
- "How do you reproduce an old result after the provider retires the model?" — You can't re-generate; that is why raw outputs are stored. You can re-grade them and compare new models against those stored results.
- "What about open-weight models?" — More control: pin weights, inference engine version and hardware settings; still expect small differences across GPU types and batch sizes.