LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How is KL divergence used in evaluating LLMs?


Where the KL term sits in RLHFPolicywrites an answerRewardmodel scores itSubtract β ×KL to thereference modelUpdatethe policyWithout the KL term the policy drifts toward answers that fool the reward model.
The KL penalty is a leash: the model may chase reward, but only as far as it stays close to a model that already writes sensible text.

What you need to know

Text
KL(P || Q) = sum over x of P(x) × log(P(x) / Q(x))

Worked numbers

Next-token distributions over three tokens from two model versions:

Text
P (original)  = [0.7, 0.2, 0.1]Q (quantized) = [0.5, 0.3, 0.2]KL(P || Q) = 0.7·ln(1.4) + 0.2·ln(0.667) + 0.1·ln(0.5) ≈ 0.085KL(Q || P) = 0.5·ln(0.714) + 0.3·ln(1.5) + 0.2·ln(2.0) ≈ 0.092

Different values in each direction: KL is asymmetric. KL(P‖Q) heavily punishes Q for giving near-zero probability where P has mass.

Where it appears across LLM training

  1. Pretraining — minimising cross-entropy equals minimising KL from the data distribution to the model.
  2. Distillation — the student minimises KL to the teacher's softened token distribution.
  3. RLHF — the policy maximises reward minus β × KL(policy ‖ reference). Without the KL term, the model finds odd outputs that fool the reward model.
  4. DPO — no separate reward model; the loss compares the tuned and reference models' log-probabilities on preferred versus rejected answers, which builds in the same KL-style anchor.
  5. Reasoning RL — methods such as GRPO often include a KL penalty to the starting model too, though some recipes reduce or drop it.

Using KL for evaluation

  • Quantization and version checks — run the same prompts through the old and new model, compute KL per token, and look at the mean and the worst cases. Tools for local model quantization report this.
  • Drift monitoring — compare distributions of topics, intents or output lengths this week versus last month.
  • Decoding comparisons — how far a sampling setting moves outputs from the base distribution.

A real-life example

A bank wants to serve its 8B support model in 4-bit instead of 16-bit to cut GPU cost. Instead of only checking a few chats by eye, the team runs 1,000 real conversations through both versions and computes the KL divergence of the next-token distributions at every position. The average is small, but the worst 1% of positions are concentrated in Devanagari text and in numbers such as account balances.

They keep 4-bit for most of the model but leave the most sensitive layers at higher precision, then rerun: the worst-case KL drops and the Hindi answers match the 16-bit model again. KL gave them a precise, per-token view that accuracy on a small test set would have missed.

Follow-up questions to expect

  • "Why is KL not a distance?" — It is asymmetric and does not satisfy the triangle inequality. Jensen–Shannon divergence is a symmetric alternative.
  • "What does β control in RLHF?" — The trade-off: small β lets the model chase reward and drift; large β keeps it close to the reference and limits improvement.
  • "Forward or reverse KL in distillation?" — Forward KL(teacher‖student) makes the student cover all of the teacher's options; reverse KL makes it focus on the main mode. LLM distillation research uses both.