Course Content
LLM Evaluation
6 sections · 50 lessons
What are reference-based vs reference-free evaluation metrics?
What you need to know
Exact match and token F1, worked
The simplest reference-based metrics come from question answering (SQuAD). Both normalise the text first: lowercase, remove punctuation and the articles "a", "an", "the".
1import re, string2from collections import Counter34def normalise(text: str) -> list[str]:5 text = text.lower()6 text = "".join(ch for ch in text if ch not in string.punctuation)7 text = re.sub(r"\b(a|an|the)\b", " ", text)8 return text.split()910def token_f1(pred: str, gold: str) -> float:11 p, g = normalise(pred), normalise(gold)12 overlap = sum((Counter(p) & Counter(g)).values())13 if overlap == 0:14 return 0.015 precision, recall = overlap / len(p), overlap / len(g)16 return 2 * precision * recall / (precision + recall)For the question "How many days of paternity leave do I get?" with gold answer "15 days":
| Prediction | Exact match | Token F1 |
|---|---|---|
| "15 days." | 1 | 1.00 |
| "15 working days" | 0 | 0.80 |
| "10 working days" | 0 | 0.40 |
"15 working days" is correct but gets EM 0; F1 gives partial credit (2 shared tokens: precision 2/3, recall 2/2, F1 0.80). But "10 working days" — a wrong answer — still earns 0.40, because "days" overlaps. F1 is fair to wording, blind to which token matters.
Reference-free metrics
- Faithfulness — are the answer's claims supported by the retrieved context?
- Answer relevance — does the answer address the question?
- Format and policy checks — valid JSON, no personal data, length limits.
- Execution — tests pass, query runs.
- Rubric judge — an LLM grades against criteria without a gold answer.
| Reference-based | Reference-free | |
|---|---|---|
| Needs labels | Yes, expensive and they age | No |
| Open-ended tasks | Penalises valid alternatives | Handles them |
| Runs on live traffic | No | Yes |
| Trust | High, if references are good | Depends on the evaluator's accuracy |
Note that faithfulness can be perfect while the answer is wrong — if the retriever fetched an outdated policy, a faithful answer repeats the outdated fact. That is why you need both kinds.
A real-life example
The HR assistant's CI suite has 200 questions with gold answers and gold source documents. It measures answer correctness (reference-based, via a judge that compares key facts with the gold answer) and context recall. In production, with no gold answers, the team samples 3% of conversations and runs faithfulness and answer relevance (reference-free).
In March, production faithfulness stayed at 0.93 while employees started complaining. The golden set, updated for the new leave policy, showed correctness dropping from 88% to 71%: the index still contained last year's policy PDF. The answers were faithful — to the wrong document. Reference-free monitoring could not see that; the reference-based set did.
Follow-up questions to expect
- "Is an LLM judge reference-based or reference-free?" — Either: you can give it a gold answer to compare against (reference-guided) or only the input and context (reference-free).
- "When is exact match the right metric?" — Closed answers: labels, IDs, numbers, dates, yes/no. Normalise first and consider accepting a list of valid aliases.
- "How do you validate a reference-free metric?" — Label a few hundred outputs by hand and measure how well the metric agrees, including its precision and recall on failures.