LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is perplexity, and what does it tell you about a model?


Probability given to each true next token0.500.250.800.100123NLL 0.22NLL 2.30, thebiggest surpriseMean NLL 1.151, so perplexity is exp(1.151), about 3.16.
One badly predicted token dominates the average, and perplexity 3.16 means roughly three equally likely choices per step.

What you need to know

A language model gives a probability to every possible next token. When you feed it a real text, you can read off the probability it gave to each token that actually came next.

Text
perplexity = exp( average over tokens of -log p(token | previous tokens) )

A worked example

Suppose a model reads a four-token text and gave the true tokens probabilities 0.5, 0.25, 0.8 and 0.1:

Python
import math# probability the model gave to each actual next tokenprobs = [0.5, 0.25, 0.8, 0.1]nll = [-math.log(p) for p in probs]          # surprise per tokenmean_nll = sum(nll) / len(nll)               # cross-entropyprint(f"mean NLL = {mean_nll:.3f}, perplexity = {math.exp(mean_nll):.2f}")
Text
mean NLL = 1.151, perplexity = 3.16

The per-token surprises are 0.69, 1.39, 0.22 and 2.30. The one token the model gave only 10% contributes most. A perplexity of 3.16 means the model was, on average, as unsure as picking among about three equally likely tokens.

With Hugging Face Transformers, model(input_ids, labels=input_ids).loss returns the mean NLL (the library shifts the labels internally), and torch.exp(loss) is the perplexity.

What it is good for

  • Tracking pretraining and fine-tuning progress on held-out text.
  • Comparing checkpoints of the same model.
  • Checking that quantisation, pruning or distillation did not damage the model.
  • Detecting distribution shift: in-domain text suddenly getting higher perplexity.

Its limits

  • Tokenizer-dependent. Different vocabularies split text into different numbers of tokens, so per-token scores are not comparable across model families. Bits per byte or per character normalise this.
  • Text-dependent. Only compare on the same evaluation text.
  • Not quality. A model can be fluent and confidently wrong.

A real-life example

A company runs an open-weight model on its own servers for the HR assistant and wants to switch to a 4-bit quantised version to halve GPU memory. They measure perplexity on 2,000 held-out paragraphs of their HR policy documents: the full-precision model scores 6.1 and the quantised one 6.4. A small rise is expected; a jump to, say, 9 would signal real damage.

But perplexity alone does not approve the switch. They also run their 200-question answer eval: 84% correct for both versions. Perplexity told them nothing broke at the language level; the task eval told them the product still works.

Follow-up questions to expect

  • "Can you compare perplexity of a Llama model and a Qwen model?" — Not directly; they use different tokenizers. Use bits per byte on the same text, or compare them on tasks.
  • "What's the relation between perplexity and cross-entropy loss?" — Perplexity is exp of the cross-entropy (in nats); minimising one minimises the other.
  • "Why not use perplexity to rank chat models?" — Because it measures prediction of a given text, not the usefulness of the model's own answers.