Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What is top-K and top-P sampling?
What you need to know
Temperature reshapes the distribution. Top-k and top-p cut off its long tail — the thousands of tokens with tiny probabilities that, together, can still get sampled and produce nonsense.
Top-k
Sort the tokens by probability, keep the top k, set the rest to zero, and renormalise so the kept ones add up to 1. k = 1 is greedy decoding.
Its weakness is that k is fixed. When the model is sure (one token at 95%), k = 40 still lets 39 bad options in. When it is genuinely unsure (50 plausible tokens), k = 40 cuts good ones.
Top-p (nucleus sampling)
Sort by probability and keep adding tokens until their total reaches p. Only those tokens can be sampled.
Next token after "Delivered in 2 working ..."days 0.82 | business 0.09 | day 0.04 | weeks 0.02 | hours 0.01 | ...top_p = 0.9 -> keep {days, business} (0.82 + 0.09 = 0.91)top_k = 4 -> keep {days, business, day, weeks}With top-p the pool shrank to two because the model was confident. Top-k kept "weeks", which would produce "2 working weeks" — a wrong promise on an e-commerce page.
Order of operations
Most libraries apply temperature, then top-k, then top-p, then sample. Details vary by library, which is another reason to change one setting at a time.
Min-p
Open-source inference engines such as llama.cpp and vLLM also offer min-p: keep tokens whose probability is at least a fraction of the top token's. Like top-p, it adapts to confidence, and it copes better with high temperatures.
Availability in 2026
| Where | Top-p | Top-k |
|---|---|---|
| Open-weight models (vLLM, llama.cpp) | Yes | Yes |
| Many hosted non-reasoning models | Yes | Some |
| Several current reasoning models | No | No |
On reasoning models the provider fixes sampling internally, and you steer with effort, the prompt and structured outputs.
A real-life example
A team runs a small open-weight model on its own GPUs to write short product blurbs. With temperature 1.0 and no cut-off, about 1 blurb in 200 contains a strange token — a random word in another language, or a broken unit like "5 Lt r". Setting top_p = 0.9 removes the long tail and the problem drops to about 1 in 5,000, while the blurbs stay varied. They do not also lower temperature, so they can attribute the fix to one change.
Follow-up questions to expect
- "What does top_p = 1.0 mean?" — No cut-off: every token stays in the pool, so temperature alone controls randomness.
- "Why is top-p preferred over top-k?" — It adapts to how confident the model is at each step, where top-k keeps the same number of tokens regardless.
- "How is this different from beam search?" — Beam search keeps several full candidate sequences and picks the highest-scoring one; it is deterministic and common in translation, but gives bland text for open-ended generation.