Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How would you set generation parameters such as temperature and top_p for two different features in the same product?


What you need to know

At each step the model produces a score (a logit) for every possible next token. Softmax turns those scores into probabilities, and the sampler picks one token. The parameters change how that pick happens.

Both control randomness, so change one and leave the other at its default. Moving both at once makes it impossible to tell which caused a change. Some newer APIs reject setting both together.

Defaults by feature

FeatureSettingWhy
Extraction, classification, SQL, tool argumentstemperature=0Creativity is a defect here
Support answers, RAGtemperature 0.2 to 0.3Natural wording; facts come from the retrieved context
Brainstorming, marketing copy variantstemperature 0.8 to 1.0, or n=5 samplesYou want diversity on purpose

The other parameters

  • max_tokens is a hard cut-off, not a target. Set it too low and you get truncated JSON that fails to parse. Always check the finish reason; "length" means it was cut.
  • stop sequences end generation cleanly at a marker.
  • frequency_penalty and presence_penalty, where an API offers them, reduce repetition. They are mostly a patch for a weak prompt.
  • seed, where offered, gives best-effort reproducibility, not a guarantee.

Reasoning models change the rules

Many reasoning models, which think before answering, fix or ignore sampling parameters. Some reject temperature entirely, and the main control becomes a reasoning-effort or thinking-budget setting. Check the model's documentation rather than copying settings across models.

Why temperature 0 is not deterministic

Even greedy decoding can differ between identical requests. On shared inference servers, your request is batched with others, and the GPU kernels are not batch-invariant: the order of floating-point additions can change with batch size, so results change in the last digits. When two tokens are nearly tied, that tiny change flips the choice, and everything after it differs. Mixture-of-experts routing can add more variation. If you need identical output for identical input, cache the response.

Python
EXTRACT = dict(temperature=0, max_tokens=800)SUPPORT = dict(temperature=0.3, max_tokens=600)IDEAS   = dict(temperature=0.9, max_tokens=400)reply = llm.complete(prompt, **SUPPORT)if reply.finish_reason == "length":    log.warning("answer truncated", prompt_id=prompt.id)

Keeping settings as named presets per feature makes them reviewable, and the finish-reason check catches truncation before it reaches a parser.

A real-life example

Scenario (illustrative numbers). A real-estate portal has two AI features: extracting fields from property listings, and writing three alternative listing descriptions for agents. Both ran at a single global temperature of 0.7.

Extraction disagreed with itself on 9% of listings when run twice, and 2% of outputs were truncated JSON because max_tokens was 300. The description writer, meanwhile, often gave three near-identical variants. The team set extraction to temperature 0 with max_tokens=800 and a finish-reason check: disagreement fell below 1% and truncation to zero. The writer moved to temperature 0.9 with three samples, and agents' "regenerate" clicks dropped by about half.

Follow-up questions to expect

  • "Is top_k the same as top_p?" — No. top_k keeps a fixed number of tokens; top_p keeps however many tokens reach the probability mass, which adapts to how confident the model is.
  • "Why not always use temperature 0?" — For open-ended text it can produce flat, repetitive wording, and for creative features you want variety.
  • "How would you get more reliable answers on a reasoning question?" — Sample several answers at moderate temperature and take the majority (self-consistency), or use a reasoning model with a higher effort setting.