Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
BLEU, ROUGE, METEOR, BERTScore, and Beyond
Here are two candidate translations of a French sentence, alongside the reference a professional translator produced.
Reference: The cat is sitting on the matCandidate A: The cat sat on the matCandidate B: On the the the the the matAny human ranks A far above B. Candidate A is a perfectly good translation that happens to use a different tense; candidate B is gibberish. Now score them with plain unigram precision — the fraction of the candidate's words that appear in the reference. Candidate A gets 5 of 6. Candidate B gets 7 of 7, a perfect score, because every one of its words does appear somewhere in the reference.
That failure is not hypothetical; it is the exact failure that shaped the design of BLEU in 2002, and the patches invented to fix it — clipped counts, higher-order n-grams, a brevity penalty — are still the reason BLEU looks the way it does. Every metric in this lesson is best understood as a specific answer to a specific way that a simpler metric got fooled. Learn the failure and the formula becomes obvious.
BLEU: precision over n-grams, with two patches
BLEU (Bilingual Evaluation Understudy) measures precision: of the n-grams the candidate produced, how many appear in the reference? Two modifications make it survive contact with a candidate like B.
Patch one: clipped counts
An n-gram can only be credited as many times as it appears in the reference. Candidate B contains "the" five times, but the reference contains it twice, so the count is clipped at 2. With "on" and "mat" credited once each, B's unigram precision collapses from 7/7 to 4/7, and the perfect score is gone.
Patch two: the brevity penalty
Precision alone rewards saying less. A candidate consisting only of the word "The" would have unigram precision 1.0. BLEU therefore multiplies by a penalty when the candidate is shorter than the reference, where c is candidate length and r is reference length:
Note what this is not: it is not a recall term. BLEU has no recall. The brevity penalty is a crude length-matching hack that discourages the degenerate short output, and it is the reason BLEU behaves oddly when your task legitimately produces outputs of varying length.
The full score
with N=4 and uniform weights wn=1/4 by default. The exponential-of-log-sum is just a geometric mean of the four precisions — and a geometric mean has a property that trips people up constantly, as the worked example shows.
Worked example, computed properly
Reference: the cat is sitting on the mat (7 tokens). Candidate: the cat sat on the mat (6 tokens).
| n | Candidate n-grams | Matched (clipped) | pn |
|---|---|---|---|
| 1 | the, cat, sat, on, the, mat (6) | the×2, cat, on, mat = 5 | 5/6 = 0.8333 |
| 2 | the cat, cat sat, sat on, on the, the mat (5) | the cat, on the, the mat = 3 | 3/5 = 0.6000 |
| 3 | the cat sat, cat sat on, sat on the, on the mat (4) | on the mat = 1 | 1/4 = 0.2500 |
| 4 | the cat sat on, cat sat on the, sat on the mat (3) | none = 0 | 0/3 = 0.0000 |
Brevity penalty: c=6<r=7, so BP=e1−7/6=e−0.1667=0.8465.
Now the geometric mean. One of the four precisions is zero, and log0=−∞, so:
Zero. A translation a human would call good scores exactly zero, because a single missing 4-gram annihilates the product. This is the most important practical fact about BLEU: on individual sentences it is close to useless, because short sentences routinely have no matching 4-gram.
BLEU is a corpus-level metric wearing the costume of a sentence-level one. Sentence-BLEU without smoothing is not a weaker signal — it is frequently the wrong signal.
Two standard rescues. First, drop to BLEU-2:
Second, apply add-one smoothing to the higher orders (add 1 to numerator and denominator for n≥2), giving p2=4/6=0.6667, p3=2/5=0.4, p4=1/4=0.25:
0.411 is a defensible score for this pair. Zero was not. If your pipeline computes sentence-level BLEU, check which smoothing method it uses — the answers differ by a factor of two or more and the library default is not always what you want.
1from sacrebleu.metrics import BLEU23# Corpus level -- this is what BLEU was designed for4bleu = BLEU()5hyps = ["The cat sat on the mat"]6# One inner list per reference *stream*, each with one entry per hypothesis7refs = [["The cat is sitting on the mat"], ["A cat sits on the mat"]]8print(bleu.corpus_score(hyps, refs)) # BLEU = 37.99 ...910# Sentence level -- needs smoothing, and say so in your report11sent = BLEU(effective_order=True)12print(sent.sentence_score("The cat sat on the mat",13 ["The cat is sitting on the mat"]))The shape of refs is the classic sacrebleu mistake. Writing both references in one inner list, [["ref 1", "ref 2"]], does not raise an error: sacrebleu reads it as one reference stream, pairs "ref 1" with your single hypothesis, silently ignores "ref 2", and prints 32.16.
Use sacrebleu rather than hand-rolled or NLTK BLEU for anything you will publish or compare across teams. BLEU scores are notoriously incomparable across implementations because tokenisation differs; sacrebleu exists specifically to pin that down and emits a signature string recording the exact configuration.
ROUGE: the same idea, pointed at recall
Summarisation has the opposite failure mode from translation. A summary that omits the main finding is bad even if every word it does contain is impeccable. So ROUGE (Recall-Oriented Understudy for Gisting Evaluation) asks: of the n-grams in the reference, how many did the candidate manage to cover?
Modern usage reports the F1 combination of ROUGE precision and recall, because pure recall is trivially gamed by outputting the whole source document.
Worked example
Reference: The company reported record profits in the third quarter (9 tokens). Candidate: The company announced record profits this quarter (7 tokens).
Matching unigrams (clipped): the, company, record, profits, quarter = 5.
The variants, and what each one is for
| Variant | Counts | Catches | Blind to |
|---|---|---|---|
| ROUGE-1 | Unigram overlap | Content-word coverage | Word order entirely |
| ROUGE-2 | Bigram overlap | Local fluency and phrasing | Long-range structure |
| ROUGE-L | Longest common subsequence | Sentence-level order without requiring contiguity | Meaning; still surface-level |
| ROUGE-Lsum | LCS computed per sentence, then unioned | Multi-sentence summaries | Same as ROUGE-L |
| ROUGE-S / SU | Skip-bigrams (pairs in order, gaps allowed) | Flexible word order | Explodes combinatorially; rarely used now |
| ROUGE-W | Weighted LCS favouring contiguous runs | Rewards long verbatim spans | Paraphrase |
ROUGE-L deserves a moment because it is the one that behaves differently from ROUGE-1 in an instructive way. Take a second candidate that scrambles the order: Record profits were announced by the company this quarter (9 tokens).
Unigram overlap is still 5 words, and now P=R=5/9=0.5556, so ROUGE-1 F1 = 0.5556 — the same recall as the well-ordered candidate. But the longest common subsequence with the reference is only 4 tokens long: record, profits, the, quarter (the the is the one in "the third quarter"). You cannot also pick up the company, which appears before record profits in the reference but after it here. So:
ROUGE-1 says 0.556, ROUGE-L says 0.444. That 11-point spread is the metric telling you the content is there but the structure is not. Reporting only ROUGE-1 hides exactly the class of problem — jumbled, reordered, incoherently assembled summaries — that ROUGE-L was invented to expose.
1from rouge_score import rouge_scorer23scorer = rouge_scorer.RougeScorer(4 ["rouge1", "rouge2", "rougeL", "rougeLsum"], use_stemmer=True5)6scores = scorer.score(7 "The company reported record profits in the third quarter",8 "Record profits were announced by the company this quarter",9)10for k, v in scores.items():11 print(f"{k:10s} P={v.precision:.3f} R={v.recall:.3f} F={v.fmeasure:.3f}")use_stemmer=True matters more than people expect: without it, report and reported are unrelated strings. Always state whether stemming was on, because it moves scores by several points.
METEOR: alignment, synonyms, and a fragmentation penalty
BLEU and ROUGE both treat words as opaque strings. Sat and sitting share no credit; rapid and quick share no credit. METEOR fixes this by building an explicit alignment between candidate and reference words in stages:
- Exact match on surface form.
- Stem match via Porter stemming (reported ↔ reporting).
- Synonym match via WordNet synsets (announced ↔ reported).
- Paraphrase match via a learned paraphrase table (in METEOR 1.5+).
Each stage runs only on words left unmatched by the previous one, and within a stage the alignment chosen is the one with the fewest crossing links. Then two things are computed.
A recall-weighted harmonic mean. METEOR weights recall nine times as heavily as precision, because empirically that correlates better with human judgement:
A fragmentation penalty. Group the aligned words into the smallest number of contiguous chunks that are adjacent in both sentences. One chunk means the candidate reproduced the reference's ordering perfectly; many chunks means the words are right but scattered.
Penalty=0.5(matcheschunks)3METEOR=Fmean×(1−Penalty)Worked example
Take the scrambled candidate again:
Record profits were announced by the company this quarteragainstThe company reported record profits in the third quarter. With synonym matching, announced now aligns to reported, giving 6 matches out of 9 candidate words and 9 reference words — but keep it simple and use the 5 exact matches to compare like with like: P=R=5/9=0.5556.Fmean=0.5556+9×0.555610×0.5556×0.5556=5.55563.0864=0.5556The matched words form three chunks — record profits, the company, and quarter — because they appear in a different order than in the reference:
Penalty=0.5(53)3=0.5×0.216=0.108METEOR=0.5556×(1−0.108)=0.5556×0.892=0.496METEOR lands at 0.496, between ROUGE-1's 0.556 and ROUGE-L's 0.444 — it credits the content but docks the disorder proportionally rather than all-or-nothing. That graded behaviour is why METEOR correlates better with human ratings than BLEU at the sentence level.
Python1import nltk2nltk.download("wordnet", quiet=True)3from nltk.translate.meteor_score import meteor_score45ref = "The company reported record profits in the third quarter".split()6hyp = "Record profits were announced by the company this quarter".split()7print(round(meteor_score([ref], hyp), 3)) # 0.413NLTK prints 0.413 here, not the 0.496 worked by hand, because it also runs stem and WordNet synonym matching and chooses its own alignment, which changes both the match count and the chunk count. That gap is itself the lesson: a METEOR score means little without the name of the implementation that produced it.
The cost of METEOR's cleverness is portability: WordNet and the paraphrase tables exist for a limited set of languages, and the tuned parameters (the 9:1 recall weighting, the 0.5 and the cube) were fitted on specific human-judgement corpora. Outside those, the calibration is a guess.
BERTScore: matching meanings instead of strings
All three metrics so far fail the same test. Consider:
TextReference: The film was extremely boringCandidate: The movie was incredibly dullSame meaning. BLEU-1 credits only "The" and "was" — 2 of 5. ROUGE-1 F1 is 0.4. METEOR does slightly better via WordNet but still misses incredibly/extremely depending on the synset. Every one of them under-rates a genuinely correct paraphrase, and this is not an edge case; it is the normal situation whenever a task admits many valid phrasings.
BERTScore replaces string identity with contextual embedding similarity. Push both sentences through a pretrained encoder (BERT, RoBERTa, DeBERTa), obtaining one vector per token that already encodes the token's context. Then greedily match each candidate token to its most similar reference token by cosine similarity:
RBERT=∣x∣1xi∈x∑x^j∈x^maxxi⊤x^jPBERT=∣x^∣1x^j∈x^∑xi∈xmaxxi⊤x^jFBERT=2PBERT+RBERTPBERT⋅RBERT(The vectors are L2-normalised, so the dot product is the cosine.) Optionally each token is weighted by its inverse document frequency, so that matching the counts for less than matching profits.
Reading the numbers correctly
Raw BERTScore has a badly miscalibrated range. Two unrelated sentences typically score around 0.75–0.85, not 0. This is because contextual embeddings of any two English tokens share a lot of geometry. The library therefore offers rescaling with baseline: subtract the expected similarity b of random sentence pairs and stretch:
F^=1−bFBERT−bWith a typical baseline b=0.83, a raw score of 0.92 rescales to (0.92−0.83)/(1−0.83)=0.09/0.17=0.529. The raw 0.92 sounds like near-perfect agreement; the rescaled 0.53 is the honest reading. Never compare raw and rescaled BERTScores, and always state which you reported.
Python1from bert_score import score23cands = ["The movie was incredibly dull"]4refs = ["The film was extremely boring"]56P, R, F = score(cands, refs, lang="en",7 model_type="microsoft/deberta-xlarge-mnli",8 rescale_with_baseline=True)9print(f"P={P.item():.3f} R={R.item():.3f} F={F.item():.3f}")BERTScore tells you whether the candidate means roughly what the reference means. It cannot tell you whether either of them is true.
That limitation is severe and worth stating plainly: a fluent, confident fabrication that is topically on-point scores well. So does a sentence with the polarity flipped, because "the drug is effective" and "the drug is not effective" are embedded close together. If factuality matters, BERTScore is not the tool.
The rest of the field
Metric Mechanism Built for Needs Watch out for CIDEr TF-IDF-weighted n-gram cosine across multiple references Image captioning 5+ references per item Meaningless with one reference SPICE Compares scene graphs (objects, attributes, relations) parsed from text Image captioning, semantics A parser Parser errors propagate silently BLEURT BERT fine-tuned to predict human ratings Translation, generation Nothing at inference Distribution shift from its training data COMET Encoder over source + hypothesis + reference, trained on human MT judgements Machine translation The source text too Reference-free variants are less reliable MoverScore Earth-mover's distance between embedding distributions Summarisation An encoder Slow at scale BARTScore Log-likelihood of one text given the other under BART Faithfulness-ish scoring A seq2seq model Length bias; longer text scores lower Perplexity Model's own uncertainty over held-out text Language-model training diagnostics Access to logits Not a quality metric at all The trend across that table is clear: from string matching (BLEU, ROUGE) to engineered linguistic matching (METEOR, SPICE) to embedding similarity (BERTScore, MoverScore) to learned regression on human judgement (BLEURT, COMET). Each step correlates better with humans and each step is harder to audit when it disagrees with you.
Choosing, and the mistakes people make while choosing
Situation Report Why Machine translation, corpus-level, comparing to published work sacreBLEU + COMET BLEU for comparability with the literature, COMET because it is the better metric Summarisation against reference summaries ROUGE-1/2/Lsum + BERTScore ROUGE for coverage, BERTScore for paraphrase tolerance Open-ended generation, no single right answer None of these Reference-based metrics are invalid without a meaningful reference; use rubric-based judging Code generation Execution: unit tests, pass@k Correctness is decidable — do not approximate it with text overlap Structured extraction (JSON, fields) Exact match, field-level F1, schema validity Same reason Grounded question answering Answer F1 + a faithfulness check against the retrieved passage Overlap says nothing about whether the answer is supported Quick sanity check during development ROUGE-L or BERTScore, with an interval Cheap, monotone-ish, adequate for catching gross breakage Named failure modes
- Comparing scores across papers or repos. BLEU 34.2 in one paper and 31.8 in another may reflect tokenisation, not quality. Only compare numbers produced by the same implementation with the same settings, and record the sacrebleu signature.
- Quoting a single-reference score as if it were absolute. With one reference, all these metrics have a low ceiling that has nothing to do with your model. A second human translator scores maybe 0.35 BLEU against the first. Your model scoring 0.32 is therefore not "68% wrong".
- Averaging sentence-level BLEU. BLEU is defined on corpus-aggregated counts. The mean of per-sentence BLEU is a different, worse statistic, and it is not what published corpus BLEU means.
- Reporting a difference without an interval. On 500 test sentences, BLEU differences under about 1 point are routinely within noise. Bootstrap-resample your test set 1,000 times and report the 2.5th and 97.5th percentiles of the difference — sacrebleu will do this with
--paired-bs. - Optimising against the metric. Beam search tuned to maximise BLEU produces literal, source-hugging translations. Selecting checkpoints on ROUGE produces extractive summaries that copy sentences. The metric rises, the product gets worse.
Every one of these metrics answers "how similar is this text to that text". If your actual question is "is this correct", "is this useful", or "is this safe", no amount of tuning the similarity metric will get you there.
What this means when you build an evaluation suite
Pick your metrics by asking what a bad output looks like on your task, then choosing the metric that specifically catches it. A summariser that drops the key finding needs a recall metric; a translator that hallucinates fluent extra clauses needs a precision metric; a paraphrase system needs an embedding metric because string overlap will punish exactly the behaviour you want.
Report at least two metrics with different mechanisms — one surface-level, one semantic — and treat disagreement between them as a finding rather than an annoyance. When ROUGE-1 rises and ROUGE-L falls, something is being reordered. When BERTScore rises and ROUGE falls, the model has started paraphrasing rather than copying, which may be exactly what you wanted or exactly what breaks your downstream regex. Neither number alone tells you that; the pair does.
Finally, validate the whole apparatus once against people. Take 200 outputs, have humans rank them, and compute the Spearman rank correlation between each candidate metric and the human ranking. If your chosen metric correlates at ρ=0.25, it is not measuring your task, and every decision you make on it afterwards is a decision made on noise dressed up to three decimal places. That one afternoon of work is worth more than any amount of metric-shopping.