Course Content
LLM Evaluation
6 sections · 50 lessons
How do BLEU, ROUGE, and BERTScore differ, and when should each be used?
What you need to know
BLEU
BLEU (2002) measures precision: what fraction of the candidate's 1- to 4-grams appear in the reference(s), combined by geometric mean. A brevity penalty stops a system winning by writing three safe words. It was designed for scoring a whole translation corpus; on single sentences it is noisy, and you should use a standard implementation such as sacrebleu so numbers are comparable.
ROUGE
ROUGE (2004) measures recall: how much of the reference's content the candidate covers. ROUGE-1 and ROUGE-2 count unigram and bigram overlap; ROUGE-L uses the longest common subsequence, so word order matters. Libraries usually report precision, recall and F1.
BERTScore
BERTScore embeds every token of both texts with a contextual encoder, matches each token to its most similar token on the other side by cosine similarity, and reports precision, recall and F1. "trainers" can match "shoes". Raw values sit in a narrow high band, so use rescale_with_baseline=True and never compare scores produced by different encoder models.
Why they fail on open-ended text
Here is a real run. The reference and two candidate product descriptions:
1import sacrebleu2from rouge_score import rouge_scorer34ref = "Lightweight running shoes with breathable mesh and a cushioned sole."5good = "Airy mesh trainers that feel light, with a soft padded sole for running."6wrong = "Lightweight running shoes with breathable leather and a hard sole."78scorer = rouge_scorer.RougeScorer(["rouge1", "rougeL"], use_stemmer=True)9for name, cand in [("good", good), ("wrong", wrong)]:10 bleu = sacrebleu.sentence_bleu(cand, [ref]).score11 r = scorer.score(ref, cand)12 print(f"{name:5} BLEU={bleu:5.1f} ROUGE-1={r['rouge1'].fmeasure:.2f} "13 f"ROUGE-L={r['rougeL'].fmeasure:.2f}")good BLEU= 4.1 ROUGE-1=0.43 ROUGE-L=0.26wrong BLEU= 45.0 ROUGE-1=0.80 ROUGE-L=0.80The accurate paraphrase scores 4.1 BLEU. The description that invents leather and a hard sole — two false claims a customer would return the shoes over — scores 45.0. Overlap metrics measure shared words, not truth.
| Metric | Good for | Avoid for |
|---|---|---|
| BLEU | Translation, corpus level | Single sentences, chat, creative text |
| ROUGE | Extractive-style summaries with fixed references | Abstractive or open-ended answers |
| BERTScore | Paraphrase-tolerant comparison to a reference | Checking facts, numbers and negation |
A real-life example
The e-commerce team first gated their description generator on ROUGE-L against descriptions written by copywriters. A new prompt dropped ROUGE-L from 0.41 to 0.33 and was nearly rejected. A human review of 100 pairs preferred the new descriptions 61 times, because they were written in fresh words while keeping every fact. Meanwhile, the old prompt's highest-ROUGE outputs often copied the reference's template but swapped attributes.
The team replaced ROUGE with two checks: a deterministic attribute check (every spec value appears, no material words outside the spec) and a pairwise LLM judge for readability. They kept ROUGE only as a drift alarm: a sudden large move means something changed and is worth a look.
Follow-up questions to expect
- "Why does BLEU have a brevity penalty?" — Because it is precision-based: a one-word output that appears in the reference would score perfectly. The penalty reduces the score when the candidate is shorter than the reference.
- "Can multiple references help?" — Yes, BLEU and ROUGE accept several references, which gives more credit to valid wordings, but it cannot cover every good answer for open-ended tasks.
- "Would BERTScore catch the leather error?" — Not reliably. "leather" and "mesh" are both materials and embed close together; a single swapped fact barely moves the score. Use entailment or attribute checks for facts.