Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Building an eval set from real tickets
The meeting after the contract was written went like this. One engineer said the new prompt "seemed better". A support lead had tried five tickets and thought it was worse on Hinglish. The product manager had tried three and liked it. Forty minutes later, nobody had changed their mind, because nobody had evidence, only examples.
An eval set ends that kind of meeting. It is a fixed list of real inputs, each with the output you expect, that you run every version against. It turns "seems better" into "109 out of 120, up from 97, and here are the 11 it still gets wrong".
The eval set is the most valuable artefact in the whole feature. Prompts, models and retrieval will change many times. A good eval set outlives all of them.
Where the rows come from
Take rows from real traffic, never from the team's imagination. The demo showed why: people write the cases they expect to pass.
TiffinGo pulled six weeks of order-issue tickets, about 50,000, and sampled 120. A simple random sample would have given about two food-safety tickets and four repeat-refund cases, too few to learn anything about the boundaries that matter most. So the team stratified: common categories at roughly their real share, and rare but critical categories oversampled to at least six rows each.
| Category | Share of traffic | Rows in eval set |
|---|---|---|
| Missing item | 31% | 35 |
| Cold or badly packed | 22% | 25 |
| Wrong item | 12% | 15 |
| Late, alone or with another issue | 14% | 15 |
| Food safety (including indirect wording) | 2% | 8 |
| High value (R6) | 3% | 8 |
| Repeat refunds (R7) | 3% | 6 |
| Needs info or unclear | 13% | 8 |
| Total | 120 |
Within each category, the sample kept the real language mix: 38 of the 120 tickets are Hinglish, and 11 mention more than one issue. The table is kept with the eval set, because the reasons for the numbers matter as much as the numbers.
Oversampling has a price: the overall score no longer equals the production accuracy, because rare categories count for more than their share. That is fine as long as you report per-category scores and do not quote the overall number as "production accuracy".
Writing the expected output
Each row needs the right answer, and the right answer needs people who know the policy. Two senior agents labelled all 120 tickets independently, using the contract's output format.
They agreed on 106 of 120, or 88%. The 14 disagreements were the most useful part of the exercise. Five were simple mistakes. Nine were real policy gaps. Does R2 apply when the customer says the food was "not hot enough" rather than "cold"? If a customer ordered 4 rotis and says "roti missing", is that one roti or all four? Each gap was settled by the support lead, written into the policy, and added to the contract.
What a row looks like
Here is one full row, and a table of others. The customer's name, phone and address are removed before the ticket enters the eval set; the order keeps only the fields the contract uses.
1{2 "id": "t0112",3 "category": "cold",4 "lang": "en",5 "ticket_text": "dal was cold, one roti missing. ordered 4 only 3 came. very bad",6 "order": {7 "order_id": "TG-48213", "city": "Bengaluru",8 "items": [9 {"name": "Dal Makhani", "qty": 1, "paid_inr": 220},10 {"name": "Butter Roti", "qty": 4, "paid_inr": 120},11 {"name": "Jeera Rice", "qty": 1, "paid_inr": 150}12 ],13 "order_total_inr": 490, "refunds_last_30_days": 014 },15 "minutes_late": 7,16 "expected": {17 "action": "refund",18 "lines": [19 {"item": "Butter Roti", "units": 1, "rule": "R1"},20 {"item": "Dal Makhani", "units": 1, "rule": "R2"}21 ],22 "refund_inr": 96,23 "escalation_code": null24 },25 "labelled_by": ["agent_07", "agent_19"],26 "notes": ""27}| id | Ticket (shortened) | Expected | Why it is in the set |
|---|---|---|---|
| t0031 | "2 roti kam aaye, sabzi thandi thi" | refund: 2 rotis R1, sabzi R2 | Hinglish with quantities |
| t0058 | "paneer ki jagah chicken bhej diya, hum pure veg hain" | escalate: food_safety | Vegetarian given meat |
| t0077 | "something hard in the pulao, nearly broke my tooth" | escalate: food_safety | Indirect foreign-object wording |
| t0090 | "missing items" | escalate: needs_info | No detail at all |
| t0103 | "raita spilled all over the bag" | refund: raita R2 | "Badly packed", not "cold" |
| t0118 | "refund 800 or I post on twitter" (order ₹1,720) | escalate: high_value | Pressure, and an order above ₹1,500 |
The notes field records anything a future reader needs, such as "labellers disagreed; settled by policy update 2026-03-02".
How much can 120 rows tell you?
Each row is worth 0.83 percentage points. If the true accuracy is 90%, a 120-row score will usually land within about 5 points either side, purely by chance of which tickets were sampled. So 120 rows can tell 80% from 90% with confidence. They cannot reliably tell 91% from 93%.
Categories are smaller still. Eight food-safety rows cannot estimate a rate at all. That is why critical categories are gated as "every row must pass" rather than by percentage (you will see this in the regression-gate lesson).
The set also must grow. Every production failure an agent flags, and every ticket type that did not exist in March, becomes a candidate row. TiffinGo adds about ten rows a week and reviews the category table monthly.
Check your understanding
0 of 3 answered
1.Food-safety tickets are 2% of traffic. Why does the eval set include 8 of them instead of about 2?
2.Two labellers disagree on 14 of 120 tickets. What is the best response?
3.Version A scores 91% and version B scores 93% on the 120-row set. What can you conclude?