Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What does the temperature setting do in an LLM?
What you need to know
A model generates text one token (a word or piece of a word) at a time. At each step it gives every token in its vocabulary a raw score called a logit. A function called softmax turns the logits into probabilities that add up to 1, and then one token is sampled from that distribution.
A worked example
Suppose the next word in a product description has three candidates with logits durable 2.0, sturdy 1.5 and rugged 0.5.
durable sturdy ruggedT = 0.5 0.71 0.26 0.04T = 1.0 0.55 0.33 0.12T = 2.0 0.44 0.35 0.21At 0.5 the model says "durable" almost every time. At 2.0 it says "rugged" one time in five. Across a 60-word description, those small choices compound into very different texts.
Choosing a value (on models that accept it)
| Task | Typical temperature |
|---|---|
| Classification, extraction, SQL, code | 0 to 0.2 |
| Q&A over documents, summaries | 0.2 to 0.5 |
| Marketing copy, brainstorming | 0.7 to 1.0 |
Two misconceptions
- Low temperature is not "more factual". It makes the model consistently pick its top guess, including a wrong one.
- T = 0 is not guaranteed identical output. Floating-point differences and batching on the provider's servers can change the top token when two are nearly tied.
The 2026 picture
Reasoning models changed this parameter. Several current models reject a non-default temperature (and top_p, top_k) with an error — for example recent Claude Opus and Sonnet models — and OpenAI's reasoning models accept only the default. On these, you control quality with the reasoning effort setting, and consistency with a clear prompt and structured outputs. Temperature still matters on non-reasoning models and on open-weight models you host yourself.
A real-life example
An e-commerce team writes descriptions for 2,000 kurtas. At temperature 0.2, descriptions for similar kurtas read almost the same — "This elegant cotton kurta is perfect for..." — and the SEO team flags duplicate text. At 1.0 they vary nicely, but 3% invent a fabric not in the specs.
Their fix has two parts. They keep 0.8 for variety on the non-reasoning model they use for copy, and they add a check in code that every fabric word in the output appears in the product's spec sheet. When they later trial a reasoning model that rejects temperature, they get variety a different way: two short example descriptions in contrasting styles and the instruction "vary the opening sentence; do not start with 'This'".
Follow-up questions to expect
- "Why can't I get the same output twice at temperature 0?" — Near-ties between tokens can flip because of floating-point and batching differences on the server. If you need repeatability, cache results and use a seed where the API supports one.
- "Should I tune temperature and top-p together?" — Usually no. Change one and leave the other at its default, or you cannot tell which one caused a change.
- "What replaces temperature on reasoning models?" — The effort setting controls how much the model thinks; the prompt and examples control style and variety.