Course Content
LLMs Deep Dive
10 sections · 40 lessons
What does temperature control in LLM text generation?
What you need to know
p(token) = softmax(logits / T)A worked example
Three candidate next tokens with logits 2.0, 1.0 and 0.1:
| Temperature | p(top) | p(second) | p(third) |
|---|---|---|---|
| 0.2 | 0.993 | 0.007 | 0.000 |
| 0.5 | 0.864 | 0.117 | 0.019 |
| 1.0 | 0.659 | 0.242 | 0.099 |
| 2.0 | 0.502 | 0.304 | 0.194 |
Dividing by a small T stretches the gaps between logits, so the top token dominates. Dividing by a large T shrinks the gaps, so the three become more equal. The order never changes — only how often the lower ones get picked.
Temperature 0
T = 0 would divide by zero, so APIs treat it as greedy decoding: always take the top token. Even then, results are not guaranteed identical across runs, because of batching differences, floating-point rounding on GPUs and, in mixture-of-experts models, routing differences. If you need exact reproducibility, cache outputs.
Choosing a value
| Task | Typical T | Why |
|---|---|---|
| Extraction, classification, JSON | 0–0.2 | One right answer; variety is just error |
| Code, SQL | 0–0.3 | Small deviations break syntax |
| Support replies | 0.3–0.7 | Natural wording, stable content |
| Brainstorming, marketing copy | 0.8–1.0 | Variety is the goal |
Above about 1.3, most models start to produce incoherent text.
Reasoning models in 2026
Several reasoning-model APIs reject or ignore the temperature setting and use a fixed sampling setup, because their training and thinking are tuned for it. On those models you control behaviour with the reasoning-effort setting, the prompt and output schemas instead.
A real-life example
A bank uses one LLM for two jobs. Job one extracts the amount, date and merchant from a customer's complaint into JSON. At T = 0.9 the model sometimes wrote "Rs 1,499" as 1499, sometimes "1,499", and once invented a date. At T = 0 the format became consistent — though when the complaint did not state a date, it still guessed one, because temperature does not add knowledge. The fix for that was a schema allowing null and a rule "never infer a date".
Job two writes Diwali offer messages in Hinglish for the marketing team. At T = 0.2, every draft started "Is Diwali, apne sapno ko...". At T = 0.9 the drafts varied enough to give the team real options, and a human picks the final one.
Follow-up questions to expect
- "Should I tune temperature and top-p together?" — Usually change one and keep the other at its default; both control randomness and interact in ways that are hard to reason about.
- "Does low temperature reduce hallucination?" — It reduces random wrong tokens, but if the model's top belief is wrong, T = 0 repeats that wrong answer every time. Grounding and verification reduce hallucination.
- "Why is my T = 0 output different between runs?" — Non-deterministic GPU arithmetic and batching; ties between nearly equal tokens can flip.