Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What does the temperature setting do in an LLM?


The same three logits at three temperatures0.710.260.040.550.330.120.440.350.21durablesturdyruggedT = 0.5T = 1.0T = 2.0Logits 2.0, 1.5 and 0.5 are divided by T before softmax.
Temperature never changes which word is most likely — it only changes how often the others get a turn.

What you need to know

A model generates text one token (a word or piece of a word) at a time. At each step it gives every token in its vocabulary a raw score called a logit. A function called softmax turns the logits into probabilities that add up to 1, and then one token is sampled from that distribution.

A worked example

Suppose the next word in a product description has three candidates with logits durable 2.0, sturdy 1.5 and rugged 0.5.

Text
            durable   sturdy   ruggedT = 0.5      0.71      0.26     0.04T = 1.0      0.55      0.33     0.12T = 2.0      0.44      0.35     0.21

At 0.5 the model says "durable" almost every time. At 2.0 it says "rugged" one time in five. Across a 60-word description, those small choices compound into very different texts.

Choosing a value (on models that accept it)

TaskTypical temperature
Classification, extraction, SQL, code0 to 0.2
Q&A over documents, summaries0.2 to 0.5
Marketing copy, brainstorming0.7 to 1.0

Two misconceptions

  • Low temperature is not "more factual". It makes the model consistently pick its top guess, including a wrong one.
  • T = 0 is not guaranteed identical output. Floating-point differences and batching on the provider's servers can change the top token when two are nearly tied.

The 2026 picture

Reasoning models changed this parameter. Several current models reject a non-default temperature (and top_p, top_k) with an error — for example recent Claude Opus and Sonnet models — and OpenAI's reasoning models accept only the default. On these, you control quality with the reasoning effort setting, and consistency with a clear prompt and structured outputs. Temperature still matters on non-reasoning models and on open-weight models you host yourself.

A real-life example

An e-commerce team writes descriptions for 2,000 kurtas. At temperature 0.2, descriptions for similar kurtas read almost the same — "This elegant cotton kurta is perfect for..." — and the SEO team flags duplicate text. At 1.0 they vary nicely, but 3% invent a fabric not in the specs.

Their fix has two parts. They keep 0.8 for variety on the non-reasoning model they use for copy, and they add a check in code that every fabric word in the output appears in the product's spec sheet. When they later trial a reasoning model that rejects temperature, they get variety a different way: two short example descriptions in contrasting styles and the instruction "vary the opening sentence; do not start with 'This'".

Follow-up questions to expect

  • "Why can't I get the same output twice at temperature 0?" — Near-ties between tokens can flip because of floating-point and batching differences on the server. If you need repeatability, cache results and use a seed where the API supports one.
  • "Should I tune temperature and top-p together?" — Usually no. Change one and leave the other at its default, or you cannot tell which one caused a change.
  • "What replaces temperature on reasoning models?" — The effort setting controls how much the model thinks; the prompt and examples control style and variety.