LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is the difference between top-k sampling and top-p sampling?


An uncertain step: 'Suggest a gift for my sister, maybe a ...'.20.18.15.12.10.08.070123456booktop-k 5 stops heretop-p 0.9stops hereThe first seven candidates add up to 0.90.
On an uncertain step top-p keeps seven ideas where top-k 5 keeps five; on a confident step it shrinks to one.

What you need to know

Two distributions, same settings

A confident step — after "Your OTP is valid for 10":

Text
" minutes" 0.90 | " mins" 0.05 | " seconds" 0.02 | " hours" 0.01 | ...

An uncertain step — after "Suggest a gift for my sister, maybe a":

Text
" book" 0.20 | " watch" 0.18 | " bag" 0.15 | " perfume" 0.12 | " plant" 0.10 | ... (many more)
SettingConfident step keepsUncertain step keeps
top-k = 55 tokens, including " hours" (a wrong fact)Only 5, cutting off good ideas
top-p = 0.91 token (" minutes" reaches 0.90)Many tokens, until the total reaches 0.9

Top-k is too loose when the model is sure and too tight when it is not. Top-p adjusts to the shape of the distribution.

How top-p works in code

Python
def top_p_filter(probs, p=0.9):    # probs: dict of token -> probability    ranked = sorted(probs.items(), key=lambda kv: kv[1], reverse=True)    kept, total = [], 0.0    for token, prob in ranked:        kept.append((token, prob))        total += prob        if total >= p:            break    return {t: pr / total for t, pr in kept}   # renormaliseprint(top_p_filter({"minutes": 0.90, "mins": 0.05, "seconds": 0.02, "hours": 0.01}))# {'minutes': 1.0}

Tokens are sorted by probability, added until the running total reaches p, and the kept probabilities are rescaled to sum to 1. Sampling then happens only among the kept tokens.

Other truncation methods

  • min-p — keep tokens whose probability is at least a fraction of the top token's (e.g. 0.1 × top). Popular in open-source serving, and robust at higher temperatures.
  • Combined — top-k as a hard cap (e.g. 50) plus top-p for adaptivity.

Order of operations

Most libraries apply temperature to the logits first, then top-k / top-p on the resulting probabilities, then sample. So a high temperature flattens the distribution, which makes top-p keep more tokens.

A real-life example

An e-commerce search assistant writes short product blurbs. With top-k = 40 and temperature 1.0, about 1 in 200 blurbs contained a wrong spec ("5000mAh" became "6000mAh"), because at confident steps top-k still left low-probability wrong numbers in the pool.

The team switches to top-p = 0.9 with temperature 0.7. At confident steps like battery numbers, the nucleus shrinks to one or two tokens, so wrong numbers disappear from the pool; at creative steps ("perfect for weekend treks"), variety remains. They still fill specs from the catalogue database into the prompt and check numbers in the output, because truncation lowers the rate of such errors but cannot remove a wrong belief the model holds with high probability.

Follow-up questions to expect

  • "What does top-p = 1.0 mean?" — No truncation: sample from the full distribution.
  • "What does top-k = 1 mean?" — Greedy decoding.
  • "Why not always use a small top-p?" — A very small p behaves like greedy decoding, bringing back repetition and bland text.