Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Building evaluation sets
By week 14, PolicyPal had a collection of test sets: the 25-message smoke set, the 60-question retrieval check, 40 tool tasks and the router's 300-message test. Each had been built for one lesson's problem. None of them looked at the whole system.
That gap showed up on a Tuesday. An engineer rewrote part of the answer prompt to make replies friendlier. The smoke set passed, 24 of 25. Nobody noticed that UK employees asking about their leave balance now got the India carry-forward rule mixed into the answer, because no test combined a UK employee, a balance check and a policy explanation. Three days and eleven confused employees later, the change was rolled back.
An evaluation set, or eval set, is a fixed collection of realistic inputs with a way to grade the outputs, run on every change. It is the single most valuable asset an applied-AI team builds, more durable than any prompt or model choice. This lesson builds PolicyPal's.
What goes in: 250 cases, weighted by risk
A good eval set is not a random sample of traffic. It is a sample shaped by what matters: common cases so the headline number reflects real use, and rare, high-cost cases so they are measured at all.
| Group | Cases | Where they come from | What they check |
|---|---|---|---|
| Policy questions with an answer | 130 | Real questions: 80 India, 35 UK, 15 Global | Right facts, right sources, right country |
| Not in the policy | 20 | Real questions HR confirmed are not covered | Says so instead of guessing |
| Tool tasks | 40 | Balance 15, tickets 15, accrual 10 | Right tool, right arguments, confirmation |
| Two intents in one message | 20 | Real multi-part messages | Nothing is silently dropped |
| Sensitive and out of scope | 20 | HR case records (redacted), odd requests | Handover or polite decline |
| Adversarial | 20 | Red-team attempts from Section 9 | No leak, no rule change |
Real traffic dominates, because synthetic questions tend to be cleaner and easier than what people actually type. The team also generated some questions from policy chunks with a model to cover sections nobody had asked about yet, then kept only the ones a human judged realistic. Be careful with those: questions written from a chunk tend to reuse its exact words, which flatters keyword search and makes retrieval look better than it is on real questions.
Every production failure becomes a case. The UK balance incident alone added five.
A case is a small, precise contract
Each case records the input, the context the system should have, and what a correct response must and must not contain. PolicyPal stores cases as JSON Lines in the repository, reviewed like code.
1{"id": "pol-uk-017", "group": "policy/uk",2 "employee": {"country": "UK", "grade": "L3"},3 "input": "Do I still get paid if I'm off sick for two days?",4 "expected_status": "answered",5 "gold_sources": [{"file": "sickness-absence-uk.pdf", "section": "3.1"}],6 "key_facts": ["Company sick pay covers up to 10 working days a year at full pay",7 "Absences of up to 7 days need a self-certification form"],8 "must_not": ["Statutory sick pay starts on day 4 and is the only pay"],9 "origin": "helpdesk ticket, Nov 2025", "added": "2026-02-03"}Notice the grading data is not a full "gold answer" to compare word for word. It is a list of key facts the answer must contain and must-not statements it must avoid. There are many correct ways to phrase an answer, and grading against exact wording would fail good answers. Key facts can be checked by an NLI model, a judge or a person, and they say precisely what "correct" means. The origin field matters too: six months later, someone will ask why a strange case exists.
How big is big enough?
With 250 cases, how much can you trust a score? Treat the pass rate as an estimate with an uncertainty. A simple approximation for a 95% interval is ±1.96 × √(p × (1 − p) ÷ n).
- Whole set: 88% passing on 250 cases gives about ±4 points.
- A group of 20 cases at 85% gives about ±16 points.
So the headline number can tell 88% from 80%, but not 88% from 86%. Group scores on 20 cases are directional only. This is why the sensitive route got its own 200-case set in Section 5.
There is a better way to compare two versions than comparing two totals. Run both on the same cases and look at the flips: cases that passed before and fail now, and the reverse.
1# policypal/evals/compare.py2import json34def load_run(path: str) -> dict[str, bool]:5 return {row["id"]: row["passed"] for row in json.load(open(path))}67def flips(before: dict[str, bool], after: dict[str, bool]) -> dict[str, list[str]]:8 """before and after map case id -> passed."""9 return {10 "broke": sorted(k for k in before if before[k] and not after.get(k, False)),11 "fixed": sorted(k for k in before if not before[k] and after.get(k, False)),12 }1314result = flips(load_run("runs/v2.3.json"), load_run("runs/v2.4.json"))15print(len(result["fixed"]), "fixed;", len(result["broke"]), "broke")16for case_id in result["broke"]:17 print("BROKE", case_id)A change that fixes 9 cases and breaks 1 is very different from one that fixes 15 and breaks 7, even though both add 8 passing cases, about 3 points. The list of broken cases is also the fastest route to why. For a formal test, the McNemar test uses exactly these two counts. In practice, reading every broken case is more useful than any p-value.
Keep the set alive, and keep part of it hidden
An eval set decays. Traffic changes: January brings new joiners asking about probation, and every policy update brings new questions. PolicyPal's rules:
- Add every confirmed production failure within a week.
- Refresh a tenth of the real-traffic cases each quarter with new questions.
- Never delete a case silently; retire it with a note when the policy it tests is withdrawn.
- Keep a held-out part, 70 of the 250 cases, that engineers do not look at while iterating. The other 180 are the development set.
The held-out part exists because every time you look at a failing case and adjust the prompt, you tune the system to that case. After months of this, the development score overstates how well the system handles new questions. The held-out score, checked before each release, tells you how much.
Check your understanding
0 of 3 answered
1.A prompt change moves the overall score from 87.6% to 86.8% on 250 cases. What should the team do?
2.Why does each PolicyPal case list key facts and must-not statements instead of one gold answer to match?
3.Why are 70 of the 250 cases held out from day-to-day work?