Course Content
LLMs Deep Dive
10 sections · 40 lessons
How is KL divergence used in evaluating LLMs?
What you need to know
KL(P || Q) = sum over x of P(x) × log(P(x) / Q(x))Worked numbers
Next-token distributions over three tokens from two model versions:
P (original) = [0.7, 0.2, 0.1]Q (quantized) = [0.5, 0.3, 0.2]KL(P || Q) = 0.7·ln(1.4) + 0.2·ln(0.667) + 0.1·ln(0.5) ≈ 0.085KL(Q || P) = 0.5·ln(0.714) + 0.3·ln(1.5) + 0.2·ln(2.0) ≈ 0.092Different values in each direction: KL is asymmetric. KL(P‖Q) heavily punishes Q for giving near-zero probability where P has mass.
Where it appears across LLM training
- Pretraining — minimising cross-entropy equals minimising KL from the data distribution to the model.
- Distillation — the student minimises KL to the teacher's softened token distribution.
- RLHF — the policy maximises reward minus β × KL(policy ‖ reference). Without the KL term, the model finds odd outputs that fool the reward model.
- DPO — no separate reward model; the loss compares the tuned and reference models' log-probabilities on preferred versus rejected answers, which builds in the same KL-style anchor.
- Reasoning RL — methods such as GRPO often include a KL penalty to the starting model too, though some recipes reduce or drop it.
Using KL for evaluation
- Quantization and version checks — run the same prompts through the old and new model, compute KL per token, and look at the mean and the worst cases. Tools for local model quantization report this.
- Drift monitoring — compare distributions of topics, intents or output lengths this week versus last month.
- Decoding comparisons — how far a sampling setting moves outputs from the base distribution.
A real-life example
A bank wants to serve its 8B support model in 4-bit instead of 16-bit to cut GPU cost. Instead of only checking a few chats by eye, the team runs 1,000 real conversations through both versions and computes the KL divergence of the next-token distributions at every position. The average is small, but the worst 1% of positions are concentrated in Devanagari text and in numbers such as account balances.
They keep 4-bit for most of the model but leave the most sensitive layers at higher precision, then rerun: the worst-case KL drops and the Hindi answers match the 16-bit model again. KL gave them a precise, per-token view that accuracy on a small test set would have missed.
Follow-up questions to expect
- "Why is KL not a distance?" — It is asymmetric and does not satisfy the triangle inequality. Jensen–Shannon divergence is a symmetric alternative.
- "What does β control in RLHF?" — The trade-off: small β lets the model chase reward and drift; large β keeps it close to the reference and limits improvement.
- "Forward or reverse KL in distillation?" — Forward KL(teacher‖student) makes the student cover all of the teacher's options; reverse KL makes it focus on the main mode. LLM distillation research uses both.