Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Perplexity and Its Limitations for LLMs
A team fine-tunes a 7-billion-parameter model on their internal documentation. Validation perplexity drops from 12.4 to 7.9 — a 36% improvement, the loss curve is beautiful, everyone is pleased. They deploy it behind their support chatbot and accuracy on their question-answering benchmark falls from 61% to 54%.
Nothing went wrong with the training. The model genuinely did get better at the thing perplexity measures. It is just that the thing perplexity measures is not the thing the chatbot needed to do, and the two came apart — as they routinely do once a model is past the point of being obviously broken.
Perplexity is worth understanding precisely, because it is the most useful diagnostic in language modelling and simultaneously the most over-interpreted number in the field. Understanding it means knowing exactly what it computes, and therefore exactly which questions it is silent on.
What perplexity computes
A language model assigns a probability to every possible next token given the tokens before it. Run it over a held-out text and you can ask: how much probability did it assign to what actually came next?
The average negative log-likelihood per token is the cross-entropy:
Perplexity is the exponential of that:
Equivalently, and more intuitively, it is the inverse geometric mean of the assigned probabilities:
The interpretation that makes it click
Perplexity is the effective branching factor: the number of equally-likely options the model was, on average, choosing between at each step.
- A model with perplexity 1 assigned probability 1 to every actual next token. It was never uncertain.
- A model that has learned nothing and outputs a uniform distribution over a 50,000-token vocabulary has perplexity 50,000 — every token is a 50,000-way guess.
- A model with perplexity 8 was, on average, as uncertain as if it were picking uniformly among 8 candidates.
Perplexity is a measure of surprise, not of quality. A model can be reliably unsurprised by text that is fluent, ungrounded and wrong.
Worked example, by hand
A model processes a five-token sequence and assigns these probabilities to the tokens that actually appeared:
| Position | P(wi∣w<i) | log2P | −lnP (nats) |
|---|---|---|---|
| 1 | 0.500 | −1 | 0.6931 |
| 2 | 0.250 | −2 | 1.3863 |
| 3 | 0.500 | −1 | 0.6931 |
| 4 | 0.125 | −3 | 2.0794 |
| 5 | 0.250 | −2 | 1.3863 |
| Sum | −9 | 6.2383 |
Mean log2 probability: −9/5=−1.8. So in base 2:
Or via nats: H=6.2383/5=1.2477 nats, and e1.2477=3.482. Same answer, as it must be — the base cancels. The model was behaving as though choosing among roughly three and a half options per token.
Notice how brutally the single 0.125 token punished the score. Because perplexity is a geometric mean, one confidently-wrong token costs more than several mildly-uncertain ones. Push any single token's probability towards zero and perplexity goes to infinity, which is why every implementation clamps or smooths.
Computing it on real text
Two details ruin more perplexity numbers than anything else: the stride used for long documents, and which tokens are counted.
1import torch2from transformers import AutoModelForCausalLM, AutoTokenizer34model_id = "gpt2-large"5tok = AutoTokenizer.from_pretrained(model_id)6model = AutoModelForCausalLM.from_pretrained(model_id).eval()78def perplexity(text: str, max_len: int = 1024, stride: int = 512) -> float:9 ids = tok(text, return_tensors="pt").input_ids10 nlls, n_tokens = [], 011 prev_end = 012 for begin in range(0, ids.size(1), stride):13 end = min(begin + max_len, ids.size(1))14 trg_len = end - prev_end # only score the NEW tokens15 chunk = ids[:, begin:end]16 target = chunk.clone()17 target[:, :-trg_len] = -100 # -100 == ignore in the loss18 with torch.no_grad():19 out = model(chunk, labels=target)20 # out.loss is the MEAN nll over scored tokens; re-weight before summing.21 # Labels are shifted by one inside the model, so count what was scored.22 n_scored = (target[:, 1:] != -100).sum().item()23 nlls.append(out.loss.item() * n_scored)24 n_tokens += n_scored25 prev_end = end26 if end == ids.size(1):27 break28 return float(torch.exp(torch.tensor(sum(nlls) / n_tokens)))2930print(round(perplexity("The quick brown fox jumps over the lazy dog. " * 40), 3))The two lines that matter are target[:, :-trg_len] = -100 and the re-weighting by n_scored. Without the mask you score the same tokens repeatedly with different amounts of context. Without the re-weighting you average means of unequal-sized groups, which quietly biases the result. A "perplexity" computed by chopping a document into independent 1024-token blocks with no overlap will be noticeably higher than the same model's true perplexity, because every block's first tokens are predicted with no context at all.
Where perplexity earns its keep
It is not a bad metric. It is a metric with a narrow, real domain of validity.
| Use | Why perplexity is right for it |
|---|---|
| Monitoring pretraining | It is the training objective. If it stops falling, something is wrong — data pipeline, learning rate, numerical instability. |
| Detecting catastrophic forgetting | Fine-tune on domain data, then measure perplexity on general text. A jump from 12 to 40 means the model has lost general capability. |
| Domain-fit diagnosis | High perplexity on your corpus versus a general corpus tells you the model has not seen text like yours. |
| Data quality auditing | Documents with extreme perplexity under a decent model are usually garbage: encoding errors, boilerplate, machine-translated spam. |
| Quantisation and compression checks | Comparing a 4-bit and a 16-bit copy of the same model on the same tokeniser is the one setting where a small perplexity delta is meaningful. |
| Membership and contamination signals | Suspiciously low perplexity on a benchmark's text suggests it was in the training data. |
Every row shares a property: the comparison is between two states of the same model family on the same tokenised data, and the question is diagnostic rather than "which model is better for users".
Seven ways perplexity misleads
1. It is not comparable across tokenisers
This is the error that produces the most confidently wrong leaderboards, so work it out fully.
Two models score the same 4,200-character document.
| Model X | Model Y | |
|---|---|---|
| Tokens for the document | 1,000 | 800 |
| Mean NLL per token | 2.1 nats | 2.4 nats |
| Perplexity | e2.1=8.17 | e2.4=11.02 |
| Total NLL for the document | 1000×2.1=2100 nats | 800×2.4=1920 nats |
By perplexity, X wins comfortably — 8.17 against 11.02. By total negative log-likelihood, Y assigned higher probability to the identical document: 1,920 nats of surprise against 2,100. Y is the better model of this text; its tokeniser simply packs more characters per token, so its per-token average is spread over fewer, harder decisions.
The fix is to normalise by something tokeniser-independent — bits per character:
Y at 0.660 bits per character is the better model, confirming the total-NLL reading and reversing the perplexity ranking entirely.
Comparing perplexity across models with different tokenisers is not slightly unfair. It can invert the true ordering. Use bits per character or bits per byte instead.
2. It rewards predictability, and correct diversity is unpredictable
Ask a model to write an opening line for a story. "It was a dark and stormy night" is high-probability and therefore low-perplexity. An original, striking opening is low-probability and therefore high-perplexity. On any open-ended task where many outputs are valid, the model that spreads probability across the good options is penalised relative to one that concentrates on the single most clichéd one. This is why models selected purely on perplexity produce bland text.
3. It says nothing about truth
Perplexity scores the model against text you supply. It never asks whether the model's own generations are true. A model can have excellent perplexity on Wikipedia and still assert that the Eiffel Tower was completed in 1912, because at generation time nothing constrains it to the text it was scored on.
4. It says nothing about safety, refusal, or instruction-following
There is no term in the formula that could. A model that ignores the system prompt entirely and one that follows it exactly can have identical perplexity on a corpus of news articles.
5. Instruction tuning makes it worse while making the model better
This is the most counter-intuitive one and the most important for modern work. Take a base model and apply supervised fine-tuning and RLHF. The resulting model is dramatically more useful — it answers questions, follows formats, refuses harmful requests. Its perplexity on raw web text typically goes up, sometimes substantially, because it now puts probability mass on assistant-style continuations that raw web text does not contain. If you selected checkpoints on validation perplexity over generic text, you would reject every aligned model in favour of the base model.
6. It is extremely sensitive to the evaluation corpus
The same model, same tokeniser, on different held-out sets:
| Evaluation corpus | Perplexity | What the number reflects |
|---|---|---|
| Wikipedia | ~11 | Heavily represented in pretraining |
| News articles | ~15 | Common but more varied |
| Reddit comments | ~24 | Noisy, informal, high entropy |
| Legal contracts | ~6 | Formulaic — low perplexity, not high competence |
| Fresh scientific abstracts | ~30 | Specialised vocabulary, post-cutoff content |
A perplexity of 6 on contracts is not evidence that the model understands contract law. Boilerplate is easy to predict. Quoting a perplexity without naming the corpus is like quoting a temperature without naming the city.
7. It gives no credit for long-range coherence
The formula is a sum of independent per-token terms. A model can predict each token well and still contradict itself between paragraph two and paragraph nine. Because most tokens are locally-determined function words and morphology, the handful of tokens that carry global consistency are a rounding error in the average.
Perplexity noise: how big a difference is real?
Perplexity is quoted to two or three decimal places and treated as exact. It is a sample statistic like any other.
Suppose you evaluate on 100,000 tokens, obtaining mean NLL 2.100 nats with a per-token standard deviation of 2.8 nats. If tokens were independent:
The 95% interval on mean NLL is 2.100±1.96×0.00886=[2.0826, 2.1174], which exponentiates to a perplexity interval of [e2.0826, e2.1174]=[8.03, 8.31].
So a competing model at 8.25 is not distinguishable from this one, despite the numbers differing in the first decimal place.
And that interval is optimistic, because tokens within a document are strongly correlated — an unusual topic makes every token in the document harder. The honest procedure is a block bootstrap over documents: resample whole documents with replacement 1,000 times, recompute perplexity each time, and take the 2.5th and 97.5th percentiles. On a 200-document corpus that interval is routinely two to three times wider than the naive per-token one.
1import numpy as np23def bootstrap_ppl(doc_nll_sums, doc_token_counts, iters=1000, seed=0):4 """Block bootstrap over documents. Returns (ppl, lo95, hi95)."""5 rng = np.random.default_rng(seed)6 nll = np.asarray(doc_nll_sums, dtype=float)7 cnt = np.asarray(doc_token_counts, dtype=float)8 point = np.exp(nll.sum() / cnt.sum())9 n = len(nll)10 draws = []11 for _ in range(iters):12 idx = rng.integers(0, n, n)13 draws.append(np.exp(nll[idx].sum() / cnt[idx].sum()))14 lo, hi = np.percentile(draws, [2.5, 97.5])15 return point, lo, hiWhat to measure instead, and alongside
| Question you actually have | Wrong tool | Right tool |
|---|---|---|
| Is model B better for our users? | Perplexity | Task accuracy or human preference, with confidence intervals |
| Does it follow instructions? | Perplexity | Constraint-satisfaction checks (format, length, required fields) |
| Is it factually reliable? | Perplexity | Claim-level verification against sources; TruthfulQA-style probes |
| Is the generated text any good? | Perplexity of the generation | Rubric-based human or LLM judging |
| Did quantisation hurt it? | Task benchmarks only | Perplexity plus task benchmarks — here perplexity is genuinely sensitive |
| Is the training run healthy? | Downstream benchmarks (too slow, too noisy) | Perplexity — it is the right tool for this |
| Was this benchmark in the training data? | Intuition | Perplexity on the benchmark text versus comparable unseen text |
The degenerate case worth knowing
Never compute perplexity of a model's own generated output as a quality signal. A model is by construction unsurprised by what it just produced — greedy decoding yields text with perplexity near 1. Repetitive output ("the the the the") also has very low perplexity. Both of these are the worst outputs the model can produce, and perplexity ranks them best. This mistake appears surprisingly often in home-grown evaluation harnesses, usually labelled "fluency".
Low perplexity on your own output measures self-consistency, not quality — and its global minimum is a model repeating one word forever.
What this means when you build an evaluation harness
Keep perplexity, but move it out of the results table and into the diagnostics panel. It belongs next to gradient norms and throughput, not next to accuracy and win rate. Give it a fixed, frozen, documented corpus so that the number means the same thing in March as it did in January, and report it as bits per character so that a tokeniser change does not silently rewrite history.
Then set an explicit expectation about what it is allowed to decide. A useful rule: perplexity may veto but never elect. If a fine-tune sends general-domain perplexity from 12 to 45, that is a genuine alarm about catastrophic forgetting and the checkpoint should be blocked. If a fine-tune moves perplexity from 8.17 to 8.10, that is inside the noise band you computed above and it decides nothing at all; the release still turns on the task benchmarks and the human evaluation.
The team in the opening had this backwards. They let a 36% perplexity improvement authorise a deployment, when the only thing that improvement licensed was the statement "the model has learned the surface statistics of our documentation". Whether it could answer a customer's question was a separate empirical question that nobody asked until the accuracy report came back seven points down.