Course Content
LLM Evaluation
6 sections · 50 lessons
How do you design a regression testing system for LLM applications?
What you need to know
What goes in the repo
Prompts, the model ID (pinned to a dated version, not a floating alias), decoding parameters, tool definitions, retrieval settings, and the golden set. A change to any of them triggers the suite.
Layered assertions
A tool like promptfoo expresses this in a config file. Each test case has variables and assertions:
1# promptfooconfig.yaml (excerpt)2tests:3 - vars:4 question: "Refund policy for damaged items?"5 assert:6 - type: is-json7 value: {type: object, required: [answer, citations]}8 - type: contains9 value: "30 days"10 - type: llm-rubric11 value: "Cites at least one policy document and does not invent a deadline."is-json with a schema and contains are deterministic. llm-rubric calls a grader model. Running promptfoo eval in CI fails the job when assertions fail. DeepEval (pytest-style tests with metrics), LangSmith experiments and the OpenAI Evals API can play the same role.
Gating rules
- Hard gates — schema validity, safety refusals, and critical cases pass at 100%.
- Soft gates — a scored metric (accuracy, judge pass rate, faithfulness) must not drop more than a chosen margin below the baseline from
main. - Per-case diff — a PR comment listing cases that went pass to fail and fail to pass, not only the average.
- Tiered schedule — 30 to 50 smoke cases on every commit; the full suite nightly and before release.
Handling non-determinism
Even at temperature 0, outputs vary slightly between runs, and many reasoning models do not accept a temperature setting at all. So run each case 2 or 3 times, and set the margin from measured noise: run the unchanged main five times and see how much the score moves on its own. A margin smaller than that noise gives random red builds.
A real-life example
A team runs a code-review bot that comments on pull requests. Its golden set has 120 past PRs with a known, seeded bug each, and 40 clean PRs. The CI gate:
- output must be valid review JSON — 100%;
- bug-catch recall on the 120 buggy PRs must stay within 4 points of
main(the unchanged system moved 3 points between runs); - false comments on the 40 clean PRs must not exceed the baseline by more than 0.2 per PR.
A prompt change that made the bot "more thorough" raised recall from 71% to 74% but doubled false comments on clean PRs. The per-case diff showed it now flagged every TODO. The gate blocked the merge; developers were spared a noisy bot.
Follow-up questions to expect
- "How do you keep CI cost down?" — Cheap checks first, a small smoke set per commit, cache results by a hash of prompt, model and input, and use a small judge model for simple binary checks.
- "What if the eval is flaky?" — Measure the noise, use repeats, widen the margin to match, and require a regression to show up on two runs before blocking.
- "How do you update the baseline?" — Only when a change merges to
main; the new scores become the baseline for the next PR.