Course Content
LLM Evaluation
6 sections · 50 lessons
How should you design benchmarks to test factual reliability?
What you need to know
What goes into each item
- The question.
- Acceptable answers (aliases, formats): "15 days", "fifteen days", "15 working days".
- The atomic facts the answer must contain.
- The source document and span, so the item can be re-checked when policies change.
- The category: trigger type, and whether it is answerable.
Three-way scoring
Grade each answer as correct, incorrect, or not attempted (abstained). OpenAI's SimpleQA benchmark (2024) uses this scheme for short fact-seeking questions. Then compare:
| Model | Correct | Incorrect | Abstained | Accuracy only |
|---|---|---|---|---|
| A | 70% | 25% | 5% | 70% |
| B | 62% | 8% | 30% | 62% |
On accuracy alone, A wins. But A is wrong on 25% of questions and B on 8%. For an HR or banking assistant, where a wrong answer is worse than "please check with HR", B is better. Make this explicit with a penalty:
score = correct - penalty * incorrectpenalty 2: A = 0.70 - 0.50 = 0.20 B = 0.62 - 0.16 = 0.46Choose the penalty from the business cost of a wrong answer versus an abstention.
Separate knowledge from grounding
Run the same questions twice: without retrieval (tests the model's memory) and with retrieval (tests grounding). If the model fails both, it is a knowledge or retrieval problem; if it fails only with retrieval, the context is misleading it.
Keep it trustworthy
- Private and refreshed — public factuality sets leak into training data. Keep a private slice; refresh time-sensitive items.
- Validate the grader — check the automatic grader against human labels.
- Size it — on 200 items, a 3-point difference is inside noise (interval about ±6 points around 70%). Use a few hundred items per slice you care about.
A real-life example
The bank builds a 400-item factuality set for its customer assistant: 100 product facts (interest rates, fees, limits) with source spans; 80 false premises ("How do I close my account through the ATM?"); 80 unanswerable questions (future rate changes); 80 numeric questions; 60 regulatory questions answered from the bank's own published policies.
Two candidate models are tested. Model X: 81% correct overall, but it abstains on only 10% of unanswerable items. Model Y: 76% correct, abstains on 72% of unanswerable items, and has a third of X's incorrect rate. With a penalty of 3 (a wrong fee quote can lead to a complaint to the banking ombudsman), Y wins clearly. The report shows all three columns per category, so the decision is visible, not hidden in one number.
Follow-up questions to expect
- "How do you stop the model from abstaining on everything?" — Include answerable items and report correct rate on them; a model that abstains everywhere scores zero correct.
- "How do you grade free-text answers automatically?" — Extract the key fact and compare with acceptable answers, or use an LLM grader with the gold answer, calibrated against human labels.
- "How often do you update the set?" — Whenever the source facts change (rates, policies), plus a scheduled review; stale gold answers make correct systems look wrong.