Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Two users ask the exact same question simultaneously. The LLM returns two different answers — one of them wrong. How do you make LLM responses deterministic — or at least consistent — for identical inputs?
What you need to know
Why temperature 0 is not enough
Temperature 0 means "always pick the most likely token". In theory that is deterministic. In production it is not, for a specific reason: your request is batched with other users' requests, and many GPU kernels are not batch-invariant. The order in which floating-point numbers are added changes with batch size, and floating-point addition is not exactly associative, so the logits change in the last decimal places. When two candidate tokens are nearly tied, that tiny change flips the choice, and everything generated after it differs. Mixture-of-experts routing can add more variation.
seed parameters, where offered, are documented as best-effort. Treat exact reproducibility from a shared API as unavailable.
The practical stack
- Cache — key on a hash of the normalised prompt plus model, version and parameters; serve repeats from Redis with a TTL. Add single-flight so two simultaneous identical questions trigger one inference.
- Pin and zero —
temperature=0and an exact, dated model version. A floating alias like "latest" is the most common cause of "it changed and nobody deployed anything". - Constrain — for classification and extraction, schema-constrained output shrinks the space of possible answers.
- Vote when correctness dominates — sample 3 to 5 answers and take the majority (self-consistency). It multiplies cost but improves accuracy on reasoning-heavy questions, and disagreement flags ambiguous inputs.
1def cached_answer(prompt: str) -> str:2 key = sha256(json.dumps({"m": MODEL_VERSION, "p": normalise(prompt), "t": 0},3 sort_keys=True).encode()).hexdigest()4 if (hit := redis.get(key)) is not None:5 return hit6 with redis.lock(f"lock:{key}", timeout=30): # single-flight7 if (hit := redis.get(key)) is not None: # someone else just computed it8 return hit9 answer = llm.complete(prompt, model=MODEL_VERSION, temperature=0)10 redis.set(key, answer, ex=86400)11 return answerThe second get inside the lock is what makes it single-flight: the waiting request finds the first one's answer.
When to aim for what
| Use case | Target | Main tools |
|---|---|---|
| Classification, extraction, routing | Same label every time | Temperature 0, schema, cache |
| FAQ and policy answers | Same facts, wording may vary | Cache, grounding, pinned version |
| Reasoning with one right answer | Right answer, reliably | Self-consistency voting |
| Creative writing | Variety is the point | None; don't force it |
Measure it
Run a fixed set of 50 queries 20 times each, nightly. Score agreement on the extracted answer or label, not string equality. A sudden drop is often your first signal that a provider changed something.
A real-life example
Scenario (illustrative numbers). A stock-broking app's assistant answers "What is the lock-in period for ELSS funds?" At market open, dozens of users ask similar questions within seconds. Support finds that 1 in 30 answers says "one year" instead of three.
The team adds a normalised-prompt cache with single-flight, pins the model version, and grounds the answer in the fund-rules document. For this FAQ-type traffic, 62% of requests become cache hits, and the nightly 20-repeat test shows factual agreement rising from 94% to 99.8%. Cost per 1,000 questions falls by more than half. The remaining variation is in wording only.
Follow-up questions to expect
- "Won't caching serve stale answers?" — Use a TTL, and include the knowledge-base version in the cache key so a document update invalidates old answers.
- "What about questions that are the same but worded differently?" — A semantic cache matches by embedding similarity; use it carefully, with a tight threshold, since near-identical questions can need different answers.
- "Can you get true determinism by self-hosting?" — Closer, with batch-invariant kernels or batch size one, but at a real throughput cost; most teams choose caching instead.