Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Two users ask the exact same question simultaneously. The LLM returns two different answers — one of them wrong. How do you make LLM responses deterministic — or at least consistent — for identical inputs?


What you need to know

Why temperature 0 is not enough

Temperature 0 means "always pick the most likely token". In theory that is deterministic. In production it is not, for a specific reason: your request is batched with other users' requests, and many GPU kernels are not batch-invariant. The order in which floating-point numbers are added changes with batch size, and floating-point addition is not exactly associative, so the logits change in the last decimal places. When two candidate tokens are nearly tied, that tiny change flips the choice, and everything generated after it differs. Mixture-of-experts routing can add more variation.

seed parameters, where offered, are documented as best-effort. Treat exact reproducibility from a shared API as unavailable.

The practical stack

  1. Cache — key on a hash of the normalised prompt plus model, version and parameters; serve repeats from Redis with a TTL. Add single-flight so two simultaneous identical questions trigger one inference.
  2. Pin and zero — temperature=0 and an exact, dated model version. A floating alias like "latest" is the most common cause of "it changed and nobody deployed anything".
  3. Constrain — for classification and extraction, schema-constrained output shrinks the space of possible answers.
  4. Vote when correctness dominates — sample 3 to 5 answers and take the majority (self-consistency). It multiplies cost but improves accuracy on reasoning-heavy questions, and disagreement flags ambiguous inputs.
Python
def cached_answer(prompt: str) -> str:    key = sha256(json.dumps({"m": MODEL_VERSION, "p": normalise(prompt), "t": 0},                            sort_keys=True).encode()).hexdigest()    if (hit := redis.get(key)) is not None:        return hit    with redis.lock(f"lock:{key}", timeout=30):       # single-flight        if (hit := redis.get(key)) is not None:        # someone else just computed it            return hit        answer = llm.complete(prompt, model=MODEL_VERSION, temperature=0)        redis.set(key, answer, ex=86400)        return answer

The second get inside the lock is what makes it single-flight: the waiting request finds the first one's answer.

When to aim for what

Use caseTargetMain tools
Classification, extraction, routingSame label every timeTemperature 0, schema, cache
FAQ and policy answersSame facts, wording may varyCache, grounding, pinned version
Reasoning with one right answerRight answer, reliablySelf-consistency voting
Creative writingVariety is the pointNone; don't force it

Measure it

Run a fixed set of 50 queries 20 times each, nightly. Score agreement on the extracted answer or label, not string equality. A sudden drop is often your first signal that a provider changed something.

A real-life example

Scenario (illustrative numbers). A stock-broking app's assistant answers "What is the lock-in period for ELSS funds?" At market open, dozens of users ask similar questions within seconds. Support finds that 1 in 30 answers says "one year" instead of three.

The team adds a normalised-prompt cache with single-flight, pins the model version, and grounds the answer in the fund-rules document. For this FAQ-type traffic, 62% of requests become cache hits, and the nightly 20-repeat test shows factual agreement rising from 94% to 99.8%. Cost per 1,000 questions falls by more than half. The remaining variation is in wording only.

Follow-up questions to expect

  • "Won't caching serve stale answers?" — Use a TTL, and include the knowledge-base version in the cache key so a document update invalidates old answers.
  • "What about questions that are the same but worded differently?" — A semantic cache matches by embedding similarity; use it carefully, with a tight threshold, since near-identical questions can need different answers.
  • "Can you get true determinism by self-hosting?" — Closer, with batch-invariant kernels or batch size one, but at a real throughput cost; most teams choose caching instead.