Course Content
LLM Evaluation
6 sections · 50 lessons
What is perplexity, and what does it tell you about a model?
What you need to know
A language model gives a probability to every possible next token. When you feed it a real text, you can read off the probability it gave to each token that actually came next.
perplexity = exp( average over tokens of -log p(token | previous tokens) )A worked example
Suppose a model reads a four-token text and gave the true tokens probabilities 0.5, 0.25, 0.8 and 0.1:
1import math23# probability the model gave to each actual next token4probs = [0.5, 0.25, 0.8, 0.1]5nll = [-math.log(p) for p in probs] # surprise per token6mean_nll = sum(nll) / len(nll) # cross-entropy7print(f"mean NLL = {mean_nll:.3f}, perplexity = {math.exp(mean_nll):.2f}")mean NLL = 1.151, perplexity = 3.16The per-token surprises are 0.69, 1.39, 0.22 and 2.30. The one token the model gave only 10% contributes most. A perplexity of 3.16 means the model was, on average, as unsure as picking among about three equally likely tokens.
With Hugging Face Transformers, model(input_ids, labels=input_ids).loss returns the mean NLL (the library shifts the labels internally), and torch.exp(loss) is the perplexity.
What it is good for
- Tracking pretraining and fine-tuning progress on held-out text.
- Comparing checkpoints of the same model.
- Checking that quantisation, pruning or distillation did not damage the model.
- Detecting distribution shift: in-domain text suddenly getting higher perplexity.
Its limits
- Tokenizer-dependent. Different vocabularies split text into different numbers of tokens, so per-token scores are not comparable across model families. Bits per byte or per character normalise this.
- Text-dependent. Only compare on the same evaluation text.
- Not quality. A model can be fluent and confidently wrong.
A real-life example
A company runs an open-weight model on its own servers for the HR assistant and wants to switch to a 4-bit quantised version to halve GPU memory. They measure perplexity on 2,000 held-out paragraphs of their HR policy documents: the full-precision model scores 6.1 and the quantised one 6.4. A small rise is expected; a jump to, say, 9 would signal real damage.
But perplexity alone does not approve the switch. They also run their 200-question answer eval: 84% correct for both versions. Perplexity told them nothing broke at the language level; the task eval told them the product still works.
Follow-up questions to expect
- "Can you compare perplexity of a Llama model and a Qwen model?" — Not directly; they use different tokenizers. Use bits per byte on the same text, or compare them on tasks.
- "What's the relation between perplexity and cross-entropy loss?" — Perplexity is exp of the cross-entropy (in nats); minimising one minimises the other.
- "Why not use perplexity to rank chat models?" — Because it measures prediction of a given text, not the usefulness of the model's own answers.