LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do BLEU, ROUGE, and BERTScore differ, and when should each be used?


Two running-shoe descriptions against one reference4.10.430.26all facts right45.00.800.80leather,hard soleBLEUROUGE-1ROUGE-Ltruthparaphrasewrong copyReal sacrebleu and rouge-score output.
Overlap metrics reward reusing the reference's words, so a description with two false claims beats an accurate paraphrase.

What you need to know

BLEU

BLEU (2002) measures precision: what fraction of the candidate's 1- to 4-grams appear in the reference(s), combined by geometric mean. A brevity penalty stops a system winning by writing three safe words. It was designed for scoring a whole translation corpus; on single sentences it is noisy, and you should use a standard implementation such as sacrebleu so numbers are comparable.

ROUGE

ROUGE (2004) measures recall: how much of the reference's content the candidate covers. ROUGE-1 and ROUGE-2 count unigram and bigram overlap; ROUGE-L uses the longest common subsequence, so word order matters. Libraries usually report precision, recall and F1.

BERTScore

BERTScore embeds every token of both texts with a contextual encoder, matches each token to its most similar token on the other side by cosine similarity, and reports precision, recall and F1. "trainers" can match "shoes". Raw values sit in a narrow high band, so use rescale_with_baseline=True and never compare scores produced by different encoder models.

Why they fail on open-ended text

Here is a real run. The reference and two candidate product descriptions:

Python
import sacrebleufrom rouge_score import rouge_scorerref = "Lightweight running shoes with breathable mesh and a cushioned sole."good = "Airy mesh trainers that feel light, with a soft padded sole for running."wrong = "Lightweight running shoes with breathable leather and a hard sole."scorer = rouge_scorer.RougeScorer(["rouge1", "rougeL"], use_stemmer=True)for name, cand in [("good", good), ("wrong", wrong)]:    bleu = sacrebleu.sentence_bleu(cand, [ref]).score    r = scorer.score(ref, cand)    print(f"{name:5} BLEU={bleu:5.1f}  ROUGE-1={r['rouge1'].fmeasure:.2f}  "          f"ROUGE-L={r['rougeL'].fmeasure:.2f}")
Text
good  BLEU=  4.1  ROUGE-1=0.43  ROUGE-L=0.26wrong BLEU= 45.0  ROUGE-1=0.80  ROUGE-L=0.80

The accurate paraphrase scores 4.1 BLEU. The description that invents leather and a hard sole — two false claims a customer would return the shoes over — scores 45.0. Overlap metrics measure shared words, not truth.

MetricGood forAvoid for
BLEUTranslation, corpus levelSingle sentences, chat, creative text
ROUGEExtractive-style summaries with fixed referencesAbstractive or open-ended answers
BERTScoreParaphrase-tolerant comparison to a referenceChecking facts, numbers and negation

A real-life example

The e-commerce team first gated their description generator on ROUGE-L against descriptions written by copywriters. A new prompt dropped ROUGE-L from 0.41 to 0.33 and was nearly rejected. A human review of 100 pairs preferred the new descriptions 61 times, because they were written in fresh words while keeping every fact. Meanwhile, the old prompt's highest-ROUGE outputs often copied the reference's template but swapped attributes.

The team replaced ROUGE with two checks: a deterministic attribute check (every spec value appears, no material words outside the spec) and a pairwise LLM judge for readability. They kept ROUGE only as a drift alarm: a sudden large move means something changed and is worth a look.

Follow-up questions to expect

  • "Why does BLEU have a brevity penalty?" — Because it is precision-based: a one-word output that appears in the reference would score perfectly. The penalty reduces the score when the candidate is shorter than the reference.
  • "Can multiple references help?" — Yes, BLEU and ROUGE accept several references, which gives more credit to valid wordings, but it cannot cover every good answer for open-ended tasks.
  • "Would BERTScore catch the leather error?" — Not reliably. "leather" and "mesh" are both materials and embed close together; a single swapped fact barely moves the score. Use entailment or attribute checks for facts.