LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you design a regression testing system for LLM applications?


What a prompt change must pass before mergePR touchesprompt,model or settingsSmoke set: 30to 50 casesHard gates:schema andsafety at 100%Soft gates: withinmeasured noise of mainPer-case diffposted on the PRRecall rose 3 points, false comments doubled, and the diff showed every TODO flagged.
An average can improve while one slice gets worse, so the gate reads per-case flips, not just the mean.

What you need to know

What goes in the repo

Prompts, the model ID (pinned to a dated version, not a floating alias), decoding parameters, tool definitions, retrieval settings, and the golden set. A change to any of them triggers the suite.

Layered assertions

A tool like promptfoo expresses this in a config file. Each test case has variables and assertions:

YAML
# promptfooconfig.yaml (excerpt)tests:  - vars:      question: "Refund policy for damaged items?"    assert:      - type: is-json        value: {type: object, required: [answer, citations]}      - type: contains        value: "30 days"      - type: llm-rubric        value: "Cites at least one policy document and does not invent a deadline."

is-json with a schema and contains are deterministic. llm-rubric calls a grader model. Running promptfoo eval in CI fails the job when assertions fail. DeepEval (pytest-style tests with metrics), LangSmith experiments and the OpenAI Evals API can play the same role.

Gating rules

  1. Hard gates — schema validity, safety refusals, and critical cases pass at 100%.
  2. Soft gates — a scored metric (accuracy, judge pass rate, faithfulness) must not drop more than a chosen margin below the baseline from main.
  3. Per-case diff — a PR comment listing cases that went pass to fail and fail to pass, not only the average.
  4. Tiered schedule — 30 to 50 smoke cases on every commit; the full suite nightly and before release.

Handling non-determinism

Even at temperature 0, outputs vary slightly between runs, and many reasoning models do not accept a temperature setting at all. So run each case 2 or 3 times, and set the margin from measured noise: run the unchanged main five times and see how much the score moves on its own. A margin smaller than that noise gives random red builds.

A real-life example

A team runs a code-review bot that comments on pull requests. Its golden set has 120 past PRs with a known, seeded bug each, and 40 clean PRs. The CI gate:

  • output must be valid review JSON — 100%;
  • bug-catch recall on the 120 buggy PRs must stay within 4 points of main (the unchanged system moved 3 points between runs);
  • false comments on the 40 clean PRs must not exceed the baseline by more than 0.2 per PR.

A prompt change that made the bot "more thorough" raised recall from 71% to 74% but doubled false comments on clean PRs. The per-case diff showed it now flagged every TODO. The gate blocked the merge; developers were spared a noisy bot.

Follow-up questions to expect

  • "How do you keep CI cost down?" — Cheap checks first, a small smoke set per commit, cache results by a hash of prompt, model and input, and use a small judge model for simple binary checks.
  • "What if the eval is flaky?" — Measure the noise, use repeats, widen the margin to match, and require a regression to show up on two runs before blocking.
  • "How do you update the baseline?" — Only when a change merges to main; the new scores become the baseline for the next PR.