Course Content
LLM Evaluation
6 sections · 50 lessons
What role do golden datasets play in model evaluation?
What you need to know
What a good row contains
- The input — the real user message, plus any context (retrieved documents, user profile) if you are testing one stage.
- The expectation — an exact label for closed tasks (
fraud); for open-ended tasks, required facts, forbidden claims and format rules, because one reference answer punishes other correct wordings. - Metadata — category, difficulty, source (production, incident, synthetic), and who labelled it.
How to build it
- Sample real traffic and group it by intent, so each intent is represented.
- Oversample the hard and rare cases. If fraud is 2% of complaints, a random 200-row set has about 4 fraud cases — too few to measure anything.
- Cover deliberately: happy path, ambiguous inputs, out-of-scope questions the system must decline, adversarial prompts, very long inputs, and every language you serve.
- Label carefully. Two people label independently; measure agreement; a third person settles disagreements. Rows people cannot agree on usually mean the criteria are unclear.
- Version it with the prompts, and log changes, so a score change can be traced to the system or to the dataset.
How big
The size sets the smallest change you can see. At 80% accuracy on 200 items, the 95% confidence interval is about ±5.5 points, so a 2-point change is invisible. Start with 50 to 100 rows, grow to several hundred for the areas that matter most.
Held-out slice
If you tune a prompt against the same 200 rows for weeks, the prompt learns those rows. Keep 20 to 30% aside and look at it only before a release — the same rule as a test set in classic ML.
A real-life example
A bank builds a complaint classifier with labels card, loan, upi, kyc, fraud and other. The golden set has 400 complaints drawn from three months of tickets. Fraud is only 2% of traffic, but the set contains 60 fraud cases, because missing fraud is the most expensive error. It includes 40 Hinglish complaints, 20 that mention two issues, and 15 that are not complaints at all.
Two analysts label every row independently. They agree on 91% of rows; most disagreements are between upi and fraud ("money debited, I never paid this merchant"). The team writes a tie-break rule into the labelling guide — unauthorised debit means fraud — and relabels. When a new model version later scores 3 points lower on fraud recall, they know it is the model, because the dataset hash did not change.
Follow-up questions to expect
- "Can you use synthetic data?" — Yes, to fill gaps such as rare intents or attacks, but mark it as synthetic and check it against real traffic; synthetic inputs are usually cleaner and easier than real ones.
- "How do you keep it current?" — Add every production incident as a row, and re-check labels when policies change; a stale label makes the eval fail correct answers.
- "Why store criteria instead of an answer?" — For open-ended outputs many answers are correct; criteria like "mentions 15 days, cites the leave policy, does not promise approval" grade them all fairly.