Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Building evaluation sets


By week 14, PolicyPal had a collection of test sets: the 25-message smoke set, the 60-question retrieval check, 40 tool tasks and the router's 300-message test. Each had been built for one lesson's problem. None of them looked at the whole system.

That gap showed up on a Tuesday. An engineer rewrote part of the answer prompt to make replies friendlier. The smoke set passed, 24 of 25. Nobody noticed that UK employees asking about their leave balance now got the India carry-forward rule mixed into the answer, because no test combined a UK employee, a balance check and a policy explanation. Three days and eleven confused employees later, the change was rolled back.

An evaluation set, or eval set, is a fixed collection of realistic inputs with a way to grade the outputs, run on every change. It is the single most valuable asset an applied-AI team builds, more durable than any prompt or model choice. This lesson builds PolicyPal's.

250 cases, weighted by risk1302040202020012345policy answersnot in policytool taskstwo intentssensitive,off-scopeadversarialAbout 4 points of uncertainty on the whole set, 16 on a group of 20.
Rare, costly cases get far more than their share of traffic, and small groups give directions, not verdicts.

What goes in: 250 cases, weighted by risk

A good eval set is not a random sample of traffic. It is a sample shaped by what matters: common cases so the headline number reflects real use, and rare, high-cost cases so they are measured at all.

GroupCasesWhere they come fromWhat they check
Policy questions with an answer130Real questions: 80 India, 35 UK, 15 GlobalRight facts, right sources, right country
Not in the policy20Real questions HR confirmed are not coveredSays so instead of guessing
Tool tasks40Balance 15, tickets 15, accrual 10Right tool, right arguments, confirmation
Two intents in one message20Real multi-part messagesNothing is silently dropped
Sensitive and out of scope20HR case records (redacted), odd requestsHandover or polite decline
Adversarial20Red-team attempts from Section 9No leak, no rule change

Real traffic dominates, because synthetic questions tend to be cleaner and easier than what people actually type. The team also generated some questions from policy chunks with a model to cover sections nobody had asked about yet, then kept only the ones a human judged realistic. Be careful with those: questions written from a chunk tend to reuse its exact words, which flatters keyword search and makes retrieval look better than it is on real questions.

Every production failure becomes a case. The UK balance incident alone added five.

A case is a small, precise contract

Each case records the input, the context the system should have, and what a correct response must and must not contain. PolicyPal stores cases as JSON Lines in the repository, reviewed like code.

JSON
{"id": "pol-uk-017", "group": "policy/uk", "employee": {"country": "UK", "grade": "L3"}, "input": "Do I still get paid if I'm off sick for two days?", "expected_status": "answered", "gold_sources": [{"file": "sickness-absence-uk.pdf", "section": "3.1"}], "key_facts": ["Company sick pay covers up to 10 working days a year at full pay",               "Absences of up to 7 days need a self-certification form"], "must_not": ["Statutory sick pay starts on day 4 and is the only pay"], "origin": "helpdesk ticket, Nov 2025", "added": "2026-02-03"}

Notice the grading data is not a full "gold answer" to compare word for word. It is a list of key facts the answer must contain and must-not statements it must avoid. There are many correct ways to phrase an answer, and grading against exact wording would fail good answers. Key facts can be checked by an NLI model, a judge or a person, and they say precisely what "correct" means. The origin field matters too: six months later, someone will ask why a strange case exists.

How big is big enough?

With 250 cases, how much can you trust a score? Treat the pass rate as an estimate with an uncertainty. A simple approximation for a 95% interval is ±1.96 × √(p × (1 − p) ÷ n).

  • Whole set: 88% passing on 250 cases gives about ±4 points.
  • A group of 20 cases at 85% gives about ±16 points.

So the headline number can tell 88% from 80%, but not 88% from 86%. Group scores on 20 cases are directional only. This is why the sensitive route got its own 200-case set in Section 5.

There is a better way to compare two versions than comparing two totals. Run both on the same cases and look at the flips: cases that passed before and fail now, and the reverse.

Python
# policypal/evals/compare.pyimport jsondef load_run(path: str) -> dict[str, bool]:    return {row["id"]: row["passed"] for row in json.load(open(path))}def flips(before: dict[str, bool], after: dict[str, bool]) -> dict[str, list[str]]:    """before and after map case id -> passed."""    return {        "broke": sorted(k for k in before if before[k] and not after.get(k, False)),        "fixed": sorted(k for k in before if not before[k] and after.get(k, False)),    }result = flips(load_run("runs/v2.3.json"), load_run("runs/v2.4.json"))print(len(result["fixed"]), "fixed;", len(result["broke"]), "broke")for case_id in result["broke"]:    print("BROKE", case_id)

A change that fixes 9 cases and breaks 1 is very different from one that fixes 15 and breaks 7, even though both add 8 passing cases, about 3 points. The list of broken cases is also the fastest route to why. For a formal test, the McNemar test uses exactly these two counts. In practice, reading every broken case is more useful than any p-value.

Keep the set alive, and keep part of it hidden

An eval set decays. Traffic changes: January brings new joiners asking about probation, and every policy update brings new questions. PolicyPal's rules:

  • Add every confirmed production failure within a week.
  • Refresh a tenth of the real-traffic cases each quarter with new questions.
  • Never delete a case silently; retire it with a note when the policy it tests is withdrawn.
  • Keep a held-out part, 70 of the 250 cases, that engineers do not look at while iterating. The other 180 are the development set.

The held-out part exists because every time you look at a failing case and adjust the prompt, you tune the system to that case. After months of this, the development score overstates how well the system handles new questions. The held-out score, checked before each release, tells you how much.

Check your understanding

0 of 3 answered

1.A prompt change moves the overall score from 87.6% to 86.8% on 250 cases. What should the team do?

2.Why does each PolicyPal case list key facts and must-not statements instead of one gold answer to match?

3.Why are 70 of the 250 cases held out from day-to-day work?