Course Content
LLM Evaluation
6 sections · 50 lessons
What is evaluation-driven development, and why is it important for AI systems?
What you need to know
In normal software, a function either returns the right value or it does not, and unit tests tell you which. An LLM feature has no such guarantee. Every change — a new sentence in the prompt, a model upgrade, a different chunk size — shifts the behaviour on many inputs at once, often in opposite directions.
The loop
- Look at real data — read 50 to 100 real inputs or traces and write down how outputs go wrong. This "error analysis" gives you failure categories.
- Write criteria — turn each category into a check: a rule, a gold answer, or a yes/no rubric question.
- Run a baseline — score the current system so you have a number to beat.
- Change one thing — prompt, model, retrieval setting or temperature.
- Re-run and compare per case — keep the change only if the target improves and nothing important regresses.
- Add new failures — every bug found later becomes a new case.
Why it matters for LLM systems
- No compiler, no type errors. A broken prompt still returns fluent text with a 200 status.
- Changes have side effects. Adding "be concise" may fix long answers and remove a required safety disclaimer.
- Models change under you. When a provider retires a model version, the eval tells you in minutes whether the replacement is safe.
- It settles arguments. "Version B passes 34 of 40 cases, version A passed 29" ends a debate that "B feels better" never will.
When the answer changes
A 30-row eval is too small to detect small improvements (see the lesson on confidence intervals), but it is enough to catch large regressions and to guide early prompt work. Grow it as the product matures.
A real-life example
A company builds a RAG assistant that answers employee questions about HR policy. Before writing a prompt, an engineer pulls 40 real questions from last quarter's HR helpdesk tickets, including 6 the assistant must decline (salary of a named colleague, legal advice). For each question she writes the required facts ("15 working days", "applies after probation") and the policy document that should be cited.
The first prompt passes 26 of 40. Adding "quote the policy line before answering" raises it to 33. A later change, "always give a helpful answer", pushes the helpful cases up but makes the assistant answer 3 of the 6 questions it should decline — so the suite drops to 32 and the change is reverted. Without the eval, the team would have shipped a policy assistant that discusses colleagues' salaries.
Follow-up questions to expect
- "How is this different from TDD?" — Same idea, but the tests are scored, not pass/fail on exact values, and they are statistical: you compare pass rates across many cases rather than expecting every single case to pass.
- "Where do the first test cases come from?" — Real logs or support tickets if the product exists; otherwise domain experts write realistic inputs, and you replace them with real traffic as soon as you have it.
- "How big should the eval be?" — Start with 30 to 50 cases you actually run on every change; grow to a few hundred once you need to detect differences of a few points.