Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 2: Non-Deterministic Model Behavior
Scenario: the same prompt, run twice through a plain Python client, comes back with different answers, and a downstream step breaks. How do you make the application behave predictably?
What you need to know
"It gave a different answer" has several possible causes. Some are in your code, and those you can remove completely.
Causes of variation, and who controls them
| Cause | Under your control? | Fix |
|---|---|---|
| Temperature above 0 | Yes | Set temperature=0 for deterministic tasks |
| Prompt built differently each run (dict ordering, timestamps, whitespace) | Yes | Load from a versioned template; normalise input |
| Floating model alias silently upgraded | Yes | Pin the dated model version |
| Batching and GPU arithmetic on the server | No | Cache; design for consistency, not identity |
Truncation because max_tokens is too low | Yes | Set it explicitly and check the finish reason |
The server-side cause is real: your request shares a batch with others, and the arithmetic inside is not batch-invariant, so near-tied tokens can flip. You cannot turn that off on a shared API. Everything else, you can.
The cache, in plain Python
1import hashlib, json23def complete(prompt: str) -> str:4 key = hashlib.sha256(json.dumps(5 {"model": MODEL, "prompt": prompt, "temperature": 0}, sort_keys=True6 ).encode()).hexdigest()7 if (hit := cache.get(key)) is not None:8 return hit9 text = client.responses.create(model=MODEL, input=prompt, temperature=0).output_text10 cache.set(key, text, ex=86400)11 return textThis uses the OpenAI Responses API; the idea is the same with any provider. sort_keys=True makes the JSON, and so the key, stable. Note that some reasoning models do not accept temperature at all; for them, the cache and a schema do the work.
Make the downstream step robust
- Constrain the shape — schema-constrained output means variation can only occur inside free-text fields, not in structure.
- Vote when it matters — for a classification or a number, sample three times and take the majority; disagreement flags an ambiguous input for review.
- Log the full context — model version, prompt version, parameters and the response, so any surprise can be reproduced.
- Measure — nightly, run a 50-prompt set 20 times and track the agreement rate.
A drop in the nightly agreement rate is often the first sign that a provider changed something.
A real-life example
Scenario (illustrative numbers). A logistics startup's script classifies delivery-exception notes into 8 categories, and a downstream job routes each note to a team. Operations complain that the same note gets routed differently on re-runs, about 1 in 12 times.
The engineer finds three causes: temperature=0.7 copied from a chat demo, a floating model alias, and a prompt that inserted the current timestamp. Fixing all three brings disagreement to 1 in 200. Adding a schema with an enum of the 8 labels and a 3-sample vote for notes where samples disagree brings it to about 1 in 2,000, and those rare cases go to a human queue. The cache means re-runs of the same note no longer call the model at all.
Follow-up questions to expect
- "Is
seedenough?" — No. Where it exists it improves repeatability, but providers describe it as best-effort; backend changes can still alter output. - "Doesn't voting triple the cost?" — Only if you vote on everything. Vote on a sample, or only when a cheap first pass has low confidence.
- "How long should the cache live?" — Until the prompt, model or source data changes; put their versions in the key so updates invalidate it automatically.