Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 2: Non-Deterministic Model Behavior


Scenario: the same prompt, run twice through a plain Python client, comes back with different answers, and a downstream step breaks. How do you make the application behave predictably?

What you need to know

"It gave a different answer" has several possible causes. Some are in your code, and those you can remove completely.

Causes of variation, and who controls them

CauseUnder your control?Fix
Temperature above 0YesSet temperature=0 for deterministic tasks
Prompt built differently each run (dict ordering, timestamps, whitespace)YesLoad from a versioned template; normalise input
Floating model alias silently upgradedYesPin the dated model version
Batching and GPU arithmetic on the serverNoCache; design for consistency, not identity
Truncation because max_tokens is too lowYesSet it explicitly and check the finish reason

The server-side cause is real: your request shares a batch with others, and the arithmetic inside is not batch-invariant, so near-tied tokens can flip. You cannot turn that off on a shared API. Everything else, you can.

The cache, in plain Python

Python
import hashlib, jsondef complete(prompt: str) -> str:    key = hashlib.sha256(json.dumps(        {"model": MODEL, "prompt": prompt, "temperature": 0}, sort_keys=True    ).encode()).hexdigest()    if (hit := cache.get(key)) is not None:        return hit    text = client.responses.create(model=MODEL, input=prompt, temperature=0).output_text    cache.set(key, text, ex=86400)    return text

This uses the OpenAI Responses API; the idea is the same with any provider. sort_keys=True makes the JSON, and so the key, stable. Note that some reasoning models do not accept temperature at all; for them, the cache and a schema do the work.

Make the downstream step robust

  1. Constrain the shape — schema-constrained output means variation can only occur inside free-text fields, not in structure.
  2. Vote when it matters — for a classification or a number, sample three times and take the majority; disagreement flags an ambiguous input for review.
  3. Log the full context — model version, prompt version, parameters and the response, so any surprise can be reproduced.
  4. Measure — nightly, run a 50-prompt set 20 times and track the agreement rate.

A drop in the nightly agreement rate is often the first sign that a provider changed something.

A real-life example

Scenario (illustrative numbers). A logistics startup's script classifies delivery-exception notes into 8 categories, and a downstream job routes each note to a team. Operations complain that the same note gets routed differently on re-runs, about 1 in 12 times.

The engineer finds three causes: temperature=0.7 copied from a chat demo, a floating model alias, and a prompt that inserted the current timestamp. Fixing all three brings disagreement to 1 in 200. Adding a schema with an enum of the 8 labels and a 3-sample vote for notes where samples disagree brings it to about 1 in 2,000, and those rare cases go to a human queue. The cache means re-runs of the same note no longer call the model at all.

Follow-up questions to expect

  • "Is seed enough?" — No. Where it exists it improves repeatability, but providers describe it as best-effort; backend changes can still alter output.
  • "Doesn't voting triple the cost?" — Only if you vote on everything. Vote on a sample, or only when a cheap first pass has low confidence.
  • "How long should the cache live?" — Until the prompt, model or source data changes; put their versions in the key so updates invalidate it automatically.