Evaluating and Testing GenAI Models

Perplexity and Its Limitations for LLMs


A team fine-tunes a 7-billion-parameter model on their internal documentation. Validation perplexity drops from 12.4 to 7.9 — a 36% improvement, the loss curve is beautiful, everyone is pleased. They deploy it behind their support chatbot and accuracy on their question-answering benchmark falls from 61% to 54%.

Nothing went wrong with the training. The model genuinely did get better at the thing perplexity measures. It is just that the thing perplexity measures is not the thing the chatbot needed to do, and the two came apart — as they routinely do once a model is past the point of being obviously broken.

Perplexity is worth understanding precisely, because it is the most useful diagnostic in language modelling and simultaneously the most over-interpreted number in the field. Understanding it means knowing exactly what it computes, and therefore exactly which questions it is silent on.

Perplexity fell 36 percent and the product got worseFine-tune oninternal docsValperplexity12.4 to 7.9Loss curvelooks beautifulShip behindthe support botQAaccuracy fallsPerplexity scores how predictable the eval corpus is, not whether answers are true.
The model learned the style of the documentation, and perplexity cannot tell that apart from learning its content.

What perplexity computes

A language model assigns a probability to every possible next token given the tokens before it. Run it over a held-out text and you can ask: how much probability did it assign to what actually came next?

The average negative log-likelihood per token is the cross-entropy:

H=−1N∑i=1Nlog⁡P(wi∣w1,…,wi−1)H = -\frac{1}{N}\sum_{i=1}^{N} \log P(w_i \mid w_1, \ldots, w_{i-1})

Perplexity is the exponential of that:

PPL=exp⁡(H)=exp⁡ ⁣(−1N∑i=1Nln⁡P(wi∣w<i))\text{PPL} = \exp(H) = \exp\!\left(-\frac{1}{N}\sum_{i=1}^{N} \ln P(w_i \mid w_{\lt i})\right)

Equivalently, and more intuitively, it is the inverse geometric mean of the assigned probabilities:

PPL=(∏i=1NP(wi∣w<i))−1/N\text{PPL} = \left(\prod_{i=1}^{N} P(w_i \mid w_{\lt i})\right)^{-1/N}

The interpretation that makes it click

Perplexity is the effective branching factor: the number of equally-likely options the model was, on average, choosing between at each step.

  • A model with perplexity 1 assigned probability 1 to every actual next token. It was never uncertain.
  • A model that has learned nothing and outputs a uniform distribution over a 50,000-token vocabulary has perplexity 50,000 — every token is a 50,000-way guess.
  • A model with perplexity 8 was, on average, as uncertain as if it were picking uniformly among 8 candidates.

Perplexity is a measure of surprise, not of quality. A model can be reliably unsurprised by text that is fluent, ungrounded and wrong.

Worked example, by hand

A model processes a five-token sequence and assigns these probabilities to the tokens that actually appeared:

PositionP(wi∣w<i)P(w_i \mid w_{\lt i})log⁡2P\log_2 P−ln⁡P-\ln P (nats)
10.500−10.6931
20.250−21.3863
30.500−10.6931
40.125−32.0794
50.250−21.3863
Sum−96.2383

Mean log2 probability: −9/5=−1.8-9/5 = -1.8. So in base 2:

PPL=21.8=3.482\text{PPL} = 2^{1.8} = 3.482

Or via nats: H=6.2383/5=1.2477H = 6.2383/5 = 1.2477 nats, and e1.2477=3.482e^{1.2477} = 3.482. Same answer, as it must be — the base cancels. The model was behaving as though choosing among roughly three and a half options per token.

Notice how brutally the single 0.125 token punished the score. Because perplexity is a geometric mean, one confidently-wrong token costs more than several mildly-uncertain ones. Push any single token's probability towards zero and perplexity goes to infinity, which is why every implementation clamps or smooths.

Computing it on real text

Two details ruin more perplexity numbers than anything else: the stride used for long documents, and which tokens are counted.

Python
import torchfrom transformers import AutoModelForCausalLM, AutoTokenizermodel_id = "gpt2-large"tok = AutoTokenizer.from_pretrained(model_id)model = AutoModelForCausalLM.from_pretrained(model_id).eval()def perplexity(text: str, max_len: int = 1024, stride: int = 512) -> float:    ids = tok(text, return_tensors="pt").input_ids    nlls, n_tokens = [], 0    prev_end = 0    for begin in range(0, ids.size(1), stride):        end = min(begin + max_len, ids.size(1))        trg_len = end - prev_end          # only score the NEW tokens        chunk = ids[:, begin:end]        target = chunk.clone()        target[:, :-trg_len] = -100       # -100 == ignore in the loss        with torch.no_grad():            out = model(chunk, labels=target)        # out.loss is the MEAN nll over scored tokens; re-weight before summing.        # Labels are shifted by one inside the model, so count what was scored.        n_scored = (target[:, 1:] != -100).sum().item()        nlls.append(out.loss.item() * n_scored)        n_tokens += n_scored        prev_end = end        if end == ids.size(1):            break    return float(torch.exp(torch.tensor(sum(nlls) / n_tokens)))print(round(perplexity("The quick brown fox jumps over the lazy dog. " * 40), 3))

The two lines that matter are target[:, :-trg_len] = -100 and the re-weighting by n_scored. Without the mask you score the same tokens repeatedly with different amounts of context. Without the re-weighting you average means of unequal-sized groups, which quietly biases the result. A "perplexity" computed by chopping a document into independent 1024-token blocks with no overlap will be noticeably higher than the same model's true perplexity, because every block's first tokens are predicted with no context at all.

Where perplexity earns its keep

It is not a bad metric. It is a metric with a narrow, real domain of validity.

UseWhy perplexity is right for it
Monitoring pretrainingIt is the training objective. If it stops falling, something is wrong — data pipeline, learning rate, numerical instability.
Detecting catastrophic forgettingFine-tune on domain data, then measure perplexity on general text. A jump from 12 to 40 means the model has lost general capability.
Domain-fit diagnosisHigh perplexity on your corpus versus a general corpus tells you the model has not seen text like yours.
Data quality auditingDocuments with extreme perplexity under a decent model are usually garbage: encoding errors, boilerplate, machine-translated spam.
Quantisation and compression checksComparing a 4-bit and a 16-bit copy of the same model on the same tokeniser is the one setting where a small perplexity delta is meaningful.
Membership and contamination signalsSuspiciously low perplexity on a benchmark's text suggests it was in the training data.

Every row shares a property: the comparison is between two states of the same model family on the same tokenised data, and the question is diagnostic rather than "which model is better for users".

Seven ways perplexity misleads

1. It is not comparable across tokenisers

This is the error that produces the most confidently wrong leaderboards, so work it out fully.

Two models score the same 4,200-character document.

Model XModel Y
Tokens for the document1,000800
Mean NLL per token2.1 nats2.4 nats
Perplexitye2.1=8.17e^{2.1} = 8.17e2.4=11.02e^{2.4} = 11.02
Total NLL for the document1000×2.1=21001000 \times 2.1 = 2100 nats800×2.4=1920800 \times 2.4 = 1920 nats

By perplexity, X wins comfortably — 8.17 against 11.02. By total negative log-likelihood, Y assigned higher probability to the identical document: 1,920 nats of surprise against 2,100. Y is the better model of this text; its tokeniser simply packs more characters per token, so its per-token average is spread over fewer, harder decisions.

The fix is to normalise by something tokeniser-independent — bits per character:

BPC=total NLL in natsln⁡2×characters\text{BPC} = \frac{\text{total NLL in nats}}{\ln 2 \times \text{characters}}

BPCX=21000.6931×4200=3029.74200=0.721BPCY=19200.6931×4200=2770.24200=0.660\text{BPC}_X = \frac{2100}{0.6931 \times 4200} = \frac{3029.7}{4200} = 0.721 \qquad \text{BPC}_Y = \frac{1920}{0.6931 \times 4200} = \frac{2770.2}{4200} = 0.660

Y at 0.660 bits per character is the better model, confirming the total-NLL reading and reversing the perplexity ranking entirely.

Comparing perplexity across models with different tokenisers is not slightly unfair. It can invert the true ordering. Use bits per character or bits per byte instead.

2. It rewards predictability, and correct diversity is unpredictable

Ask a model to write an opening line for a story. "It was a dark and stormy night" is high-probability and therefore low-perplexity. An original, striking opening is low-probability and therefore high-perplexity. On any open-ended task where many outputs are valid, the model that spreads probability across the good options is penalised relative to one that concentrates on the single most clichéd one. This is why models selected purely on perplexity produce bland text.

3. It says nothing about truth

Perplexity scores the model against text you supply. It never asks whether the model's own generations are true. A model can have excellent perplexity on Wikipedia and still assert that the Eiffel Tower was completed in 1912, because at generation time nothing constrains it to the text it was scored on.

4. It says nothing about safety, refusal, or instruction-following

There is no term in the formula that could. A model that ignores the system prompt entirely and one that follows it exactly can have identical perplexity on a corpus of news articles.

5. Instruction tuning makes it worse while making the model better

This is the most counter-intuitive one and the most important for modern work. Take a base model and apply supervised fine-tuning and RLHF. The resulting model is dramatically more useful — it answers questions, follows formats, refuses harmful requests. Its perplexity on raw web text typically goes up, sometimes substantially, because it now puts probability mass on assistant-style continuations that raw web text does not contain. If you selected checkpoints on validation perplexity over generic text, you would reject every aligned model in favour of the base model.

6. It is extremely sensitive to the evaluation corpus

The same model, same tokeniser, on different held-out sets:

Evaluation corpusPerplexityWhat the number reflects
Wikipedia~11Heavily represented in pretraining
News articles~15Common but more varied
Reddit comments~24Noisy, informal, high entropy
Legal contracts~6Formulaic — low perplexity, not high competence
Fresh scientific abstracts~30Specialised vocabulary, post-cutoff content

A perplexity of 6 on contracts is not evidence that the model understands contract law. Boilerplate is easy to predict. Quoting a perplexity without naming the corpus is like quoting a temperature without naming the city.

7. It gives no credit for long-range coherence

The formula is a sum of independent per-token terms. A model can predict each token well and still contradict itself between paragraph two and paragraph nine. Because most tokens are locally-determined function words and morphology, the handful of tokens that carry global consistency are a rounding error in the average.

Perplexity noise: how big a difference is real?

Perplexity is quoted to two or three decimal places and treated as exact. It is a sample statistic like any other.

Suppose you evaluate on 100,000 tokens, obtaining mean NLL 2.100 nats with a per-token standard deviation of 2.8 nats. If tokens were independent:

SE(Hˉ)=2.8100,000=2.8316.2=0.00886 natsSE(\bar{H}) = \frac{2.8}{\sqrt{100{,}000}} = \frac{2.8}{316.2} = 0.00886 \text{ nats}

The 95% interval on mean NLL is 2.100±1.96×0.00886=[2.0826, 2.1174]2.100 \pm 1.96 \times 0.00886 = [2.0826,\ 2.1174], which exponentiates to a perplexity interval of [e2.0826, e2.1174]=[8.03, 8.31][e^{2.0826},\ e^{2.1174}] = [8.03,\ 8.31].

So a competing model at 8.25 is not distinguishable from this one, despite the numbers differing in the first decimal place.

And that interval is optimistic, because tokens within a document are strongly correlated — an unusual topic makes every token in the document harder. The honest procedure is a block bootstrap over documents: resample whole documents with replacement 1,000 times, recompute perplexity each time, and take the 2.5th and 97.5th percentiles. On a 200-document corpus that interval is routinely two to three times wider than the naive per-token one.

Python
import numpy as npdef bootstrap_ppl(doc_nll_sums, doc_token_counts, iters=1000, seed=0):    """Block bootstrap over documents. Returns (ppl, lo95, hi95)."""    rng = np.random.default_rng(seed)    nll = np.asarray(doc_nll_sums, dtype=float)    cnt = np.asarray(doc_token_counts, dtype=float)    point = np.exp(nll.sum() / cnt.sum())    n = len(nll)    draws = []    for _ in range(iters):        idx = rng.integers(0, n, n)        draws.append(np.exp(nll[idx].sum() / cnt[idx].sum()))    lo, hi = np.percentile(draws, [2.5, 97.5])    return point, lo, hi

What to measure instead, and alongside

Question you actually haveWrong toolRight tool
Is model B better for our users?PerplexityTask accuracy or human preference, with confidence intervals
Does it follow instructions?PerplexityConstraint-satisfaction checks (format, length, required fields)
Is it factually reliable?PerplexityClaim-level verification against sources; TruthfulQA-style probes
Is the generated text any good?Perplexity of the generationRubric-based human or LLM judging
Did quantisation hurt it?Task benchmarks onlyPerplexity plus task benchmarks — here perplexity is genuinely sensitive
Is the training run healthy?Downstream benchmarks (too slow, too noisy)Perplexity — it is the right tool for this
Was this benchmark in the training data?IntuitionPerplexity on the benchmark text versus comparable unseen text

The degenerate case worth knowing

Never compute perplexity of a model's own generated output as a quality signal. A model is by construction unsurprised by what it just produced — greedy decoding yields text with perplexity near 1. Repetitive output ("the the the the") also has very low perplexity. Both of these are the worst outputs the model can produce, and perplexity ranks them best. This mistake appears surprisingly often in home-grown evaluation harnesses, usually labelled "fluency".

Low perplexity on your own output measures self-consistency, not quality — and its global minimum is a model repeating one word forever.

What this means when you build an evaluation harness

Keep perplexity, but move it out of the results table and into the diagnostics panel. It belongs next to gradient norms and throughput, not next to accuracy and win rate. Give it a fixed, frozen, documented corpus so that the number means the same thing in March as it did in January, and report it as bits per character so that a tokeniser change does not silently rewrite history.

Then set an explicit expectation about what it is allowed to decide. A useful rule: perplexity may veto but never elect. If a fine-tune sends general-domain perplexity from 12 to 45, that is a genuine alarm about catastrophic forgetting and the checkpoint should be blocked. If a fine-tune moves perplexity from 8.17 to 8.10, that is inside the noise band you computed above and it decides nothing at all; the release still turns on the task benchmarks and the human evaluation.

The team in the opening had this backwards. They let a 36% perplexity improvement authorise a deployment, when the only thing that improvement licensed was the statement "the model has learned the surface statistics of our documentation". Whether it could answer a customer's question was a separate empirical question that nobody asked until the accuracy report came back seven points down.