Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the difference between top-k sampling and top-p sampling?
What you need to know
Two distributions, same settings
A confident step — after "Your OTP is valid for 10":
" minutes" 0.90 | " mins" 0.05 | " seconds" 0.02 | " hours" 0.01 | ...An uncertain step — after "Suggest a gift for my sister, maybe a":
" book" 0.20 | " watch" 0.18 | " bag" 0.15 | " perfume" 0.12 | " plant" 0.10 | ... (many more)| Setting | Confident step keeps | Uncertain step keeps |
|---|---|---|
| top-k = 5 | 5 tokens, including " hours" (a wrong fact) | Only 5, cutting off good ideas |
| top-p = 0.9 | 1 token (" minutes" reaches 0.90) | Many tokens, until the total reaches 0.9 |
Top-k is too loose when the model is sure and too tight when it is not. Top-p adjusts to the shape of the distribution.
How top-p works in code
1def top_p_filter(probs, p=0.9):2 # probs: dict of token -> probability3 ranked = sorted(probs.items(), key=lambda kv: kv[1], reverse=True)4 kept, total = [], 0.05 for token, prob in ranked:6 kept.append((token, prob))7 total += prob8 if total >= p:9 break10 return {t: pr / total for t, pr in kept} # renormalise1112print(top_p_filter({"minutes": 0.90, "mins": 0.05, "seconds": 0.02, "hours": 0.01}))13# {'minutes': 1.0}Tokens are sorted by probability, added until the running total reaches p, and the kept probabilities are rescaled to sum to 1. Sampling then happens only among the kept tokens.
Other truncation methods
- min-p — keep tokens whose probability is at least a fraction of the top token's (e.g. 0.1 × top). Popular in open-source serving, and robust at higher temperatures.
- Combined — top-k as a hard cap (e.g. 50) plus top-p for adaptivity.
Order of operations
Most libraries apply temperature to the logits first, then top-k / top-p on the resulting probabilities, then sample. So a high temperature flattens the distribution, which makes top-p keep more tokens.
A real-life example
An e-commerce search assistant writes short product blurbs. With top-k = 40 and temperature 1.0, about 1 in 200 blurbs contained a wrong spec ("5000mAh" became "6000mAh"), because at confident steps top-k still left low-probability wrong numbers in the pool.
The team switches to top-p = 0.9 with temperature 0.7. At confident steps like battery numbers, the nucleus shrinks to one or two tokens, so wrong numbers disappear from the pool; at creative steps ("perfect for weekend treks"), variety remains. They still fill specs from the catalogue database into the prompt and check numbers in the output, because truncation lowers the rate of such errors but cannot remove a wrong belief the model holds with high probability.
Follow-up questions to expect
- "What does top-p = 1.0 mean?" — No truncation: sample from the full distribution.
- "What does top-k = 1 mean?" — Greedy decoding.
- "Why not always use a small top-p?" — A very small p behaves like greedy decoding, bringing back repetition and bland text.