LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What are reference-based vs reference-free evaluation metrics?


Gold answer: 15 days11.0000.8000.40exact matchtoken F115 days.15 working days10 working days
Exact match fails a correct paraphrase, while token F1 gives partial credit to a wrong number.

What you need to know

Exact match and token F1, worked

The simplest reference-based metrics come from question answering (SQuAD). Both normalise the text first: lowercase, remove punctuation and the articles "a", "an", "the".

Python
import re, stringfrom collections import Counterdef normalise(text: str) -> list[str]:    text = text.lower()    text = "".join(ch for ch in text if ch not in string.punctuation)    text = re.sub(r"\b(a|an|the)\b", " ", text)    return text.split()def token_f1(pred: str, gold: str) -> float:    p, g = normalise(pred), normalise(gold)    overlap = sum((Counter(p) & Counter(g)).values())    if overlap == 0:        return 0.0    precision, recall = overlap / len(p), overlap / len(g)    return 2 * precision * recall / (precision + recall)

For the question "How many days of paternity leave do I get?" with gold answer "15 days":

PredictionExact matchToken F1
"15 days."11.00
"15 working days"00.80
"10 working days"00.40

"15 working days" is correct but gets EM 0; F1 gives partial credit (2 shared tokens: precision 2/3, recall 2/2, F1 0.80). But "10 working days" — a wrong answer — still earns 0.40, because "days" overlaps. F1 is fair to wording, blind to which token matters.

Reference-free metrics

  • Faithfulness — are the answer's claims supported by the retrieved context?
  • Answer relevance — does the answer address the question?
  • Format and policy checks — valid JSON, no personal data, length limits.
  • Execution — tests pass, query runs.
  • Rubric judge — an LLM grades against criteria without a gold answer.
Reference-basedReference-free
Needs labelsYes, expensive and they ageNo
Open-ended tasksPenalises valid alternativesHandles them
Runs on live trafficNoYes
TrustHigh, if references are goodDepends on the evaluator's accuracy

Note that faithfulness can be perfect while the answer is wrong — if the retriever fetched an outdated policy, a faithful answer repeats the outdated fact. That is why you need both kinds.

A real-life example

The HR assistant's CI suite has 200 questions with gold answers and gold source documents. It measures answer correctness (reference-based, via a judge that compares key facts with the gold answer) and context recall. In production, with no gold answers, the team samples 3% of conversations and runs faithfulness and answer relevance (reference-free).

In March, production faithfulness stayed at 0.93 while employees started complaining. The golden set, updated for the new leave policy, showed correctness dropping from 88% to 71%: the index still contained last year's policy PDF. The answers were faithful — to the wrong document. Reference-free monitoring could not see that; the reference-based set did.

Follow-up questions to expect

  • "Is an LLM judge reference-based or reference-free?" — Either: you can give it a gold answer to compare against (reference-guided) or only the input and context (reference-free).
  • "When is exact match the right metric?" — Closed answers: labels, IDs, numbers, dates, yes/no. Normalise first and consider accepting a list of valid aliases.
  • "How do you validate a reference-free metric?" — Label a few hundred outputs by hand and measure how well the metric agrees, including its precision and recall on failures.