LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

Why is perplexity less useful for instruction-tuned models?


Three fine-tuning checkpoints for product descriptions8.244%7.158%6.451%perplexitypairwise win rateepoch 1epoch 2epoch 3Epoch 3 copies stock phrases from the training data.
The lowest-perplexity checkpoint is not the one people prefer, so the behavioural eval picks when to stop.

What you need to know

Three reasons perplexity misleads here

  1. Objective mismatch. A base model is trained to predict web text. After SFT and preference training, it prefers a specific style: polite, structured, cautious. That style is less typical of web text, so perplexity on a general corpus can rise while users like the model more. Optimising for perplexity would pull the model back toward raw text.
  2. Nothing single to score. Perplexity needs a fixed text. For "summarise this complaint" there are many good summaries. Scoring the likelihood of one reference summary measures agreement with one arbitrary wording.
  3. Confidence is not correctness. Preference-tuned models become sharper: they put high probability on their preferred answers, including wrong ones. A hallucinated policy detail can have very low perplexity under the model that produced it.

Where it still helps

  • Comparing base-model checkpoints during pretraining.
  • Checking that quantisation, merging or pruning did not damage the model — compare the same model before and after on the same text.
  • Domain adaptation: perplexity on your domain text going down during continued pretraining is a useful health signal.
  • Detecting unusual inputs: a sudden rise in perplexity of incoming traffic can indicate a new kind of input.

What to use instead

Behavioural evals: task success on a golden set, rubric-based judge scores, pairwise preference between versions, and for closed tasks, accuracy and F1.

A real-life example

The e-commerce team fine-tunes a small open-weight model on 20,000 approved product descriptions. They compare three checkpoints:

CheckpointPerplexity on catalogue textPairwise win rate vs current generatorAttribute check pass
Epoch 18.244%96%
Epoch 27.158%99%
Epoch 36.451%99%

Epoch 3 has the lowest perplexity but loses to epoch 2 in pairwise judging: reviewers see it copying stock phrases from the training data. The team ships epoch 2. Perplexity told them training was converging; the behavioural eval told them when to stop.

Follow-up questions to expect

  • "Could you compute perplexity of a gold answer to score a chat model?" — You can, and it is used in some research, but it only measures agreement with one reference and is sensitive to wording; it is a weak signal for open tasks.
  • "Does RLHF always increase perplexity?" — Not always, but it commonly does on general text because it moves the model away from the pretraining distribution; the point is that the direction of perplexity no longer tells you about quality.
  • "What is a likelihood-based eval that still makes sense for chat models?" — Comparing log-probabilities of fixed answer options, as in some multiple-choice benchmarks; but generation-based evals are closer to real use.