LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What does temperature control in LLM text generation?


The same three logits at three temperatures0.8640.1170.0190.6590.2420.0990.5020.3040.194logit 2.0logit 1.0logit 0.1T = 0.5T = 1.0T = 2.0
Temperature never reorders the candidates; it only decides how often the second and third get a turn.

What you need to know

Text
p(token) = softmax(logits / T)

A worked example

Three candidate next tokens with logits 2.0, 1.0 and 0.1:

Temperaturep(top)p(second)p(third)
0.20.9930.0070.000
0.50.8640.1170.019
1.00.6590.2420.099
2.00.5020.3040.194

Dividing by a small T stretches the gaps between logits, so the top token dominates. Dividing by a large T shrinks the gaps, so the three become more equal. The order never changes — only how often the lower ones get picked.

Temperature 0

T = 0 would divide by zero, so APIs treat it as greedy decoding: always take the top token. Even then, results are not guaranteed identical across runs, because of batching differences, floating-point rounding on GPUs and, in mixture-of-experts models, routing differences. If you need exact reproducibility, cache outputs.

Choosing a value

TaskTypical TWhy
Extraction, classification, JSON0–0.2One right answer; variety is just error
Code, SQL0–0.3Small deviations break syntax
Support replies0.3–0.7Natural wording, stable content
Brainstorming, marketing copy0.8–1.0Variety is the goal

Above about 1.3, most models start to produce incoherent text.

Reasoning models in 2026

Several reasoning-model APIs reject or ignore the temperature setting and use a fixed sampling setup, because their training and thinking are tuned for it. On those models you control behaviour with the reasoning-effort setting, the prompt and output schemas instead.

A real-life example

A bank uses one LLM for two jobs. Job one extracts the amount, date and merchant from a customer's complaint into JSON. At T = 0.9 the model sometimes wrote "Rs 1,499" as 1499, sometimes "1,499", and once invented a date. At T = 0 the format became consistent — though when the complaint did not state a date, it still guessed one, because temperature does not add knowledge. The fix for that was a schema allowing null and a rule "never infer a date".

Job two writes Diwali offer messages in Hinglish for the marketing team. At T = 0.2, every draft started "Is Diwali, apne sapno ko...". At T = 0.9 the drafts varied enough to give the team real options, and a human picks the final one.

Follow-up questions to expect

  • "Should I tune temperature and top-p together?" — Usually change one and keep the other at its default; both control randomness and interact in ways that are hard to reason about.
  • "Does low temperature reduce hallucination?" — It reduces random wrong tokens, but if the model's top belief is wrong, T = 0 repeats that wrong answer every time. Grounding and verification reduce hallucination.
  • "Why is my T = 0 output different between runs?" — Non-deterministic GPU arithmetic and batching; ties between nearly equal tokens can flip.