Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
How would you set generation parameters such as temperature and top_p for two different features in the same product?
What you need to know
At each step the model produces a score (a logit) for every possible next token. Softmax turns those scores into probabilities, and the sampler picks one token. The parameters change how that pick happens.
Both control randomness, so change one and leave the other at its default. Moving both at once makes it impossible to tell which caused a change. Some newer APIs reject setting both together.
Defaults by feature
| Feature | Setting | Why |
|---|---|---|
| Extraction, classification, SQL, tool arguments | temperature=0 | Creativity is a defect here |
| Support answers, RAG | temperature 0.2 to 0.3 | Natural wording; facts come from the retrieved context |
| Brainstorming, marketing copy variants | temperature 0.8 to 1.0, or n=5 samples | You want diversity on purpose |
The other parameters
max_tokensis a hard cut-off, not a target. Set it too low and you get truncated JSON that fails to parse. Always check the finish reason; "length" means it was cut.stopsequences end generation cleanly at a marker.frequency_penaltyandpresence_penalty, where an API offers them, reduce repetition. They are mostly a patch for a weak prompt.seed, where offered, gives best-effort reproducibility, not a guarantee.
Reasoning models change the rules
Many reasoning models, which think before answering, fix or ignore sampling parameters. Some reject temperature entirely, and the main control becomes a reasoning-effort or thinking-budget setting. Check the model's documentation rather than copying settings across models.
Why temperature 0 is not deterministic
Even greedy decoding can differ between identical requests. On shared inference servers, your request is batched with others, and the GPU kernels are not batch-invariant: the order of floating-point additions can change with batch size, so results change in the last digits. When two tokens are nearly tied, that tiny change flips the choice, and everything after it differs. Mixture-of-experts routing can add more variation. If you need identical output for identical input, cache the response.
1EXTRACT = dict(temperature=0, max_tokens=800)2SUPPORT = dict(temperature=0.3, max_tokens=600)3IDEAS = dict(temperature=0.9, max_tokens=400)45reply = llm.complete(prompt, **SUPPORT)6if reply.finish_reason == "length":7 log.warning("answer truncated", prompt_id=prompt.id)Keeping settings as named presets per feature makes them reviewable, and the finish-reason check catches truncation before it reaches a parser.
A real-life example
Scenario (illustrative numbers). A real-estate portal has two AI features: extracting fields from property listings, and writing three alternative listing descriptions for agents. Both ran at a single global temperature of 0.7.
Extraction disagreed with itself on 9% of listings when run twice, and 2% of outputs were truncated JSON because max_tokens was 300. The description writer, meanwhile, often gave three near-identical variants. The team set extraction to temperature 0 with max_tokens=800 and a finish-reason check: disagreement fell below 1% and truncation to zero. The writer moved to temperature 0.9 with three samples, and agents' "regenerate" clicks dropped by about half.
Follow-up questions to expect
- "Is top_k the same as top_p?" — No. top_k keeps a fixed number of tokens; top_p keeps however many tokens reach the probability mass, which adapts to how confident the model is.
- "Why not always use temperature 0?" — For open-ended text it can produce flat, repetitive wording, and for creative features you want variety.
- "How would you get more reliable answers on a reasoning question?" — Sample several answers at moderate temperature and take the majority (self-consistency), or use a reasoning model with a higher effort setting.