Course Content
LLM Evaluation
6 sections · 50 lessons
Why is perplexity less useful for instruction-tuned models?
What you need to know
Three reasons perplexity misleads here
- Objective mismatch. A base model is trained to predict web text. After SFT and preference training, it prefers a specific style: polite, structured, cautious. That style is less typical of web text, so perplexity on a general corpus can rise while users like the model more. Optimising for perplexity would pull the model back toward raw text.
- Nothing single to score. Perplexity needs a fixed text. For "summarise this complaint" there are many good summaries. Scoring the likelihood of one reference summary measures agreement with one arbitrary wording.
- Confidence is not correctness. Preference-tuned models become sharper: they put high probability on their preferred answers, including wrong ones. A hallucinated policy detail can have very low perplexity under the model that produced it.
Where it still helps
- Comparing base-model checkpoints during pretraining.
- Checking that quantisation, merging or pruning did not damage the model — compare the same model before and after on the same text.
- Domain adaptation: perplexity on your domain text going down during continued pretraining is a useful health signal.
- Detecting unusual inputs: a sudden rise in perplexity of incoming traffic can indicate a new kind of input.
What to use instead
Behavioural evals: task success on a golden set, rubric-based judge scores, pairwise preference between versions, and for closed tasks, accuracy and F1.
A real-life example
The e-commerce team fine-tunes a small open-weight model on 20,000 approved product descriptions. They compare three checkpoints:
| Checkpoint | Perplexity on catalogue text | Pairwise win rate vs current generator | Attribute check pass |
|---|---|---|---|
| Epoch 1 | 8.2 | 44% | 96% |
| Epoch 2 | 7.1 | 58% | 99% |
| Epoch 3 | 6.4 | 51% | 99% |
Epoch 3 has the lowest perplexity but loses to epoch 2 in pairwise judging: reviewers see it copying stock phrases from the training data. The team ships epoch 2. Perplexity told them training was converging; the behavioural eval told them when to stop.
Follow-up questions to expect
- "Could you compute perplexity of a gold answer to score a chat model?" — You can, and it is used in some research, but it only measures agreement with one reference and is sensitive to wording; it is a weak signal for open tasks.
- "Does RLHF always increase perplexity?" — Not always, but it commonly does on general text because it moves the model away from the pretraining distribution; the point is that the direction of perplexity no longer tells you about quality.
- "What is a likelihood-based eval that still makes sense for chat models?" — Comparing log-probabilities of fixed answer options, as in some multiple-choice benchmarks; but generation-based evals are closer to real use.