Building AI Features in Python Backends

What stays the same in your backend, and what doesn't


You already know how to call a payment gateway, a search cluster or another team's service. You set a timeout, you handle 4xx and 5xx, you retry what is safe to retry, you log the call and you write a test with a mock. Almost all of that still applies to a model call. That is good news: you do not need a new discipline, just a few adjustments.

The adjustments matter, though. ShipFast's first prototype called the model with a 30-second timeout, retried on any error, logged the full customer message at INFO level and had one test that checked the endpoint returned 200. It worked in the demo. In the first week with real traffic it leaked phone numbers into the log system, spent $140 on retries during a provider incident and returned "Reschedule" (capital R) to a database column that expected "reschedule".

This lesson lists what actually changes, so you can adjust before you ship, not after.

The same habits, a different kind of dependencyAn internal service call• 5 to 50 ms• Same input, same output• A wrong answer is a bug, fixed once• Fails with a status codeA model call• 0.5 to 10 s, set by output length• Same input, sometimes a new answer• Wrong answers are a rate you measure• Can fail with a fluent, wrong 200
Timeouts, retries, mocks and validation all still apply; what changes is that correctness becomes a rate and every call has a token price.

Six properties that differ

A typical internal service call

  • 5 to 50 ms
  • Same input gives the same output
  • Wrong answers are bugs you can fix once
  • Cost is your own servers, roughly flat
  • Fails with a status code
  • Data stays inside your network

A model call

  • 500 ms to 10 s, depending on output length
  • Same input can give different wording, or a different answer
  • Wrong answers are a rate you measure and reduce
  • Cost is per token, for every call
  • Can "fail" with a fluent, wrong 200 response
  • Your customer's text goes to a third party

Take them one at a time.

Latency. A model writes its answer one token at a time. A token is a piece of a word, roughly 4 characters of English. If a model produces 80 tokens per second and your answer is 120 tokens, generation alone takes 1.5 seconds, plus a few hundred milliseconds before the first token. Short outputs are fast; long outputs are slow. This single fact shapes ShipFast's design: classification returns about 40 tokens and is quick, the drafted reply returns about 90 and is the slowest call.

Non-determinism. Send the same message twice and you may get two differently worded drafts. Usually the classification label is stable, but on borderline messages ("parcel late again, can you deliver Monday instead?" — complaint or reschedule?) it can flip. Lesson 1.4 is about living with this.

Wrong answers as a rate. With normal code, a bug is a bug: fix it and it stays fixed. With a model, you have an error rate, say 4% of messages misclassified. You reduce it with better prompts, better validation and better models, and you measure it with an evaluation set. You never get it to zero, so the system around the model must handle the errors that remain.

Cost per call. Your database costs roughly the same whether you run 10 or 10,000 queries an hour. A model costs money for every input and output token. At ShipFast's volume, the difference between a 400-token and a 1,200-token prompt is thousands of dollars a year.

Silent failure. The most dangerous failure is a 200 OK with a confident, wrong answer. The model does not throw an exception when it misreads "don't deliver tomorrow". Your validation and your tests have to catch what the status code will not.

Data leaving your network. The customer's message, with their phone number and address, goes to the model provider. You need to know the provider's retention terms and log carefully on your side. Section 5 covers this.

What still works, and what needs adjusting

HabitStill works?Adjustment
TimeoutsYesSet them per call type: 5 s for classify, 15 s for a long draft. Never leave the SDK default of 10 minutes.
RetriesYes, carefullyRetry 429, 5xx and timeouts only, with backoff and a total deadline. Never retry a 400.
Mocks in testsYesMock the model client, not the HTTP layer. Add a separate evaluation suite that calls the real model.
LoggingYes, carefullyLog tokens, latency, cost and prompt version. Do not log raw customer text by default.
Schema validationMore than everEvery model reply is untrusted input. Validate it like a request from the internet.
Capacity planningChangedThe provider limits your requests and tokens per minute. You share that limit across all customers.

Putting a number on one call

Before you add a model call to any endpoint, estimate its latency and cost. You need three numbers: input tokens, output tokens and the price per million tokens.

For ShipFast's classifier, the instructions are about 350 tokens and a typical customer message is about 60. The reply is about 40 tokens. At a price of $5 per million input tokens and $25 per million output tokens:

  • Input: 410 × $5 / 1,000,000 = $0.00205
  • Output: 40 × $25 / 1,000,000 = $0.00100
  • Total: about $0.003 per message, or about $120 a day at 40,000 messages

For latency: about 400 ms before the first token, then 40 tokens at roughly 80 tokens per second, which is 500 ms. Call it 0.9 seconds, with a long tail when the provider is busy. Now you can decide, before writing code, whether that fits your endpoint's latency budget and your team's cost budget.

Check your understanding

0 of 3 answered

1.A model returns HTTP 200 with a well-formed reply that says a damaged-parcel message is a "complaint". What kind of failure is this?

2.Your classifier uses 400 input tokens and 40 output tokens. Prices are $1 per million input and $5 per million output. What does one call cost?

3.Which retry policy fits a model call?