Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Living with an answer that changes every time
Most of your backend is deterministic. Given the same request and the same database state, an endpoint returns the same response. Your tests depend on this, your caches depend on it, and so does your ability to debug: reproduce the input and you reproduce the bug.
A model call breaks that. Send ShipFast's classifier the message "parcel late again, can you deliver Monday instead?" ten times and you might get reschedule eight times and complaint twice. Ask for a draft reply ten times and you get ten different wordings. Nothing is broken. This is how the model works.
The question for a backend engineer is not how to make the model deterministic. It is how to put a non-deterministic part inside a system that still behaves predictably. The answer is a pattern this course uses everywhere: a deterministic shell around a non-deterministic core.
Why the same input gives different output
A model does not look up an answer. For each position in its reply, it computes a probability for every possible next token and then picks one. With a setting called temperature above zero, it samples from those probabilities, so less likely tokens are sometimes chosen. Once one token differs, everything after it can differ too.
Older guides say "set temperature to 0 for repeatable output". That advice is weaker than it sounds. Several current models no longer accept a temperature setting at all, and even where you can set it to 0, providers do not promise identical output: batching on shared hardware and tiny numeric differences can still change a token. Treat the model as non-deterministic, always, and design for it.
The deterministic shell
The shell is ordinary code that runs before and after every model call. It has four parts.
- Constrain the input — send a closed list of labels, clear rules and only the data needed. The fewer choices the model has, the less it can vary.
- Validate the output — parse it into a typed object and reject anything outside the allowed values. An invalid answer becomes a known failure, not a surprise in the database.
- Record the result — store the validated output with the message ID, prompt version and model. From then on, the answer to "what did we decide for this message?" is a database read, not a new model call.
- Fall back on failure — when the call fails or the output stays invalid, choose a safe default, such as sending the message to a human queue.
Step 3 is the one engineers most often skip, and it is the one that makes the rest of the system deterministic again. Once a result is stored, every later read returns the same answer. Your retries, your dashboards and your audit trail all see one stable decision, even though the model might say something different if you asked again.
Idempotency: never ask twice for the same message
ShipFast receives messages through a gateway that retries when it does not get a response quickly. Without care, a retried delivery triggers a second triage: two model calls, two charges and, worse, maybe two different answers, with one agent seeing reschedule and another seeing complaint for the same message.
The fix is idempotency, keyed on the message ID the gateway already sends. Here is the core of it, with a dictionary standing in for a database table.
1from shipfast.schemas import InboundMessage, TriageResult23RESULTS: dict[str, TriageResult] = {} # in production: a table with message_id as primary key456async def triage_once(msg: InboundMessage, run_triage) -> TriageResult:7 """Run the model pipeline at most once per message_id."""8 existing = RESULTS.get(msg.message_id)9 if existing is not None:10 return existing # duplicate delivery: same answer, zero cost11 result = await run_triage(msg) # the only non-deterministic step12 RESULTS[msg.message_id] = result # from now on, reads are deterministic13 return resultIn production, the table has message_id as its primary key and you insert a "pending" row before calling the model, so two workers cannot both start on the same message. Section 4 builds that version. The idea is the same: the model runs once, and its answer becomes a fact.
Testing a non-deterministic component
You cannot write assert classify(msg) == "reschedule" against a real model and expect it to pass every time. So ShipFast splits testing in two, a split you will use from Section 3 onwards.
Unit tests (every commit)
- Replace the model with a scripted fake
- Test the shell: parsing, validation, repair, routing, fallbacks
- Fast, free and fully deterministic
- Answer: "does our code handle every kind of reply?"
Evaluations (when prompts or models change)
- Call the real model on 300 labelled messages
- Measure accuracy and per-label recall
- Cost about $1 and take a few minutes
- Answer: "how often is the model right?"
Unit tests guard your code. Evaluations guard the model's behaviour. You need both, and you should never confuse them: a green unit test suite says nothing about whether the classifier is accurate.
Check your understanding
0 of 3 answered
1.The messaging gateway delivers the same message twice, 3 seconds apart. What should ShipFast do on the second delivery?
2.A teammate sets temperature to 0 and says the classifier is now deterministic, so the unit tests can call the real model. What is wrong with this?
3.Which step of the deterministic shell turns a model's answer into a stable fact the rest of the system can rely on?