LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How should you design benchmarks to test factual reliability?


Two models on the same factual test70%25%5%0.2062%8%30%0.46correctincorrectabstainedscore, penalty 2model Amodel B
Accuracy alone picks model A; counting wrong answers as costlier than abstentions picks the model that knows when to stop.

What you need to know

What goes into each item

  • The question.
  • Acceptable answers (aliases, formats): "15 days", "fifteen days", "15 working days".
  • The atomic facts the answer must contain.
  • The source document and span, so the item can be re-checked when policies change.
  • The category: trigger type, and whether it is answerable.

Three-way scoring

Grade each answer as correct, incorrect, or not attempted (abstained). OpenAI's SimpleQA benchmark (2024) uses this scheme for short fact-seeking questions. Then compare:

ModelCorrectIncorrectAbstainedAccuracy only
A70%25%5%70%
B62%8%30%62%

On accuracy alone, A wins. But A is wrong on 25% of questions and B on 8%. For an HR or banking assistant, where a wrong answer is worse than "please check with HR", B is better. Make this explicit with a penalty:

Text
score = correct - penalty * incorrectpenalty 2:  A = 0.70 - 0.50 = 0.20    B = 0.62 - 0.16 = 0.46

Choose the penalty from the business cost of a wrong answer versus an abstention.

Separate knowledge from grounding

Run the same questions twice: without retrieval (tests the model's memory) and with retrieval (tests grounding). If the model fails both, it is a knowledge or retrieval problem; if it fails only with retrieval, the context is misleading it.

Keep it trustworthy

  • Private and refreshed — public factuality sets leak into training data. Keep a private slice; refresh time-sensitive items.
  • Validate the grader — check the automatic grader against human labels.
  • Size it — on 200 items, a 3-point difference is inside noise (interval about ±6 points around 70%). Use a few hundred items per slice you care about.

A real-life example

The bank builds a 400-item factuality set for its customer assistant: 100 product facts (interest rates, fees, limits) with source spans; 80 false premises ("How do I close my account through the ATM?"); 80 unanswerable questions (future rate changes); 80 numeric questions; 60 regulatory questions answered from the bank's own published policies.

Two candidate models are tested. Model X: 81% correct overall, but it abstains on only 10% of unanswerable items. Model Y: 76% correct, abstains on 72% of unanswerable items, and has a third of X's incorrect rate. With a penalty of 3 (a wrong fee quote can lead to a complaint to the banking ombudsman), Y wins clearly. The report shows all three columns per category, so the decision is visible, not hidden in one number.

Follow-up questions to expect

  • "How do you stop the model from abstaining on everything?" — Include answerable items and report correct rate on them; a model that abstains everywhere scores zero correct.
  • "How do you grade free-text answers automatically?" — Extract the key fact and compare with acceptable answers, or use an LLM grader with the gold answer, calibrated against human labels.
  • "How often do you update the set?" — Whenever the source facts change (rates, policies), plus a scheduled review; stale gold answers make correct systems look wrong.