Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is top-K and top-P sampling?


Next token after 'Delivered in 2 working ...'.82.09.04.02.0101234daysbusinessweeks — keptby top-k 4Top-p 0.9 stops at 0.91 after two tokens; top-k 4 keeps four, including the wrong promise.
Top-p shrinks the pool when the model is confident, which is exactly when a fixed top-k lets bad tokens in.

What you need to know

Temperature reshapes the distribution. Top-k and top-p cut off its long tail — the thousands of tokens with tiny probabilities that, together, can still get sampled and produce nonsense.

Top-k

Sort the tokens by probability, keep the top k, set the rest to zero, and renormalise so the kept ones add up to 1. k = 1 is greedy decoding.

Its weakness is that k is fixed. When the model is sure (one token at 95%), k = 40 still lets 39 bad options in. When it is genuinely unsure (50 plausible tokens), k = 40 cuts good ones.

Top-p (nucleus sampling)

Sort by probability and keep adding tokens until their total reaches p. Only those tokens can be sampled.

Text
Next token after "Delivered in 2 working ..."days 0.82 | business 0.09 | day 0.04 | weeks 0.02 | hours 0.01 | ...top_p = 0.9  -> keep {days, business}        (0.82 + 0.09 = 0.91)top_k = 4    -> keep {days, business, day, weeks}

With top-p the pool shrank to two because the model was confident. Top-k kept "weeks", which would produce "2 working weeks" — a wrong promise on an e-commerce page.

Order of operations

Most libraries apply temperature, then top-k, then top-p, then sample. Details vary by library, which is another reason to change one setting at a time.

Min-p

Open-source inference engines such as llama.cpp and vLLM also offer min-p: keep tokens whose probability is at least a fraction of the top token's. Like top-p, it adapts to confidence, and it copes better with high temperatures.

Availability in 2026

WhereTop-pTop-k
Open-weight models (vLLM, llama.cpp)YesYes
Many hosted non-reasoning modelsYesSome
Several current reasoning modelsNoNo

On reasoning models the provider fixes sampling internally, and you steer with effort, the prompt and structured outputs.

A real-life example

A team runs a small open-weight model on its own GPUs to write short product blurbs. With temperature 1.0 and no cut-off, about 1 blurb in 200 contains a strange token — a random word in another language, or a broken unit like "5 Lt r". Setting top_p = 0.9 removes the long tail and the problem drops to about 1 in 5,000, while the blurbs stay varied. They do not also lower temperature, so they can attribute the fix to one change.

Follow-up questions to expect

  • "What does top_p = 1.0 mean?" — No cut-off: every token stays in the pool, so temperature alone controls randomness.
  • "Why is top-p preferred over top-k?" — It adapts to how confident the model is at each step, where top-k keeps the same number of tokens regardless.
  • "How is this different from beam search?" — Beam search keeps several full candidate sequences and picks the highest-scoring one; it is deterministic and common in translation, but gives bland text for open-ended generation.