Building AI Features in Python Backends

Anatomy of a request: messages, parameters, tokens


In the last section you saw that a model call is an HTTPS request. Now you need to know that request well enough to control it. Every field changes something you care about: what the model sees, how long it can write, how long it takes and what you pay.

ShipFast's first classifier call was copied from a blog post. It had no limit on output length, sent the instructions as part of the user's text, ignored the response's stop reason and never looked at the token counts. It worked, but nobody could say what it cost or why it sometimes returned half a sentence. This lesson takes a request apart so that does not happen to you.

What you send, and what comes backsystemmessagesmax_tokensefforttext blocksstop_reasoninput_tokensoutput_tokensRequestResponseThinking tokens count toward max_tokens and are billed as output.
max_tokens and stop_reason are a pair: one sets the ceiling, the other tells you whether the answer hit it.

The parts of a request

A request has three kinds of content and a handful of settings.

  • System prompt — your instructions: who the model is, what the task is, the allowed labels, the rules. This is written by you and is the same for every call of one feature.
  • Messages — the conversation, as a list of turns with a role of user or assistant. For ShipFast, this is usually one user turn holding the customer's message. In the repair step of Section 3, you add the model's previous reply as an assistant turn and your correction as a new user turn.
  • Model — which model runs the request. Keep it in configuration, never hard-coded in the call, because you will change it.
  • max_tokens — the most tokens the model may generate. It is a hard stop, not a target. It is also your cost ceiling for output.
  • Effort or reasoning settings — many current models can think before answering. More thinking can raise accuracy on hard tasks, but thinking tokens are generated output: they count toward max_tokens, add latency and are billed. For short tasks like ShipFast's classifier, low effort is the right starting point.

Some providers and older models also accept temperature, which controls how random the sampling is. Several current models no longer accept it. Do not build anything that depends on it.

The same call with two SDKs

Here is ShipFast's classifier request with Anthropic's Python SDK. It is synchronous to keep it short; ShipFast itself uses the async client.

Python
import anthropicclient = anthropic.Anthropic()        # reads ANTHROPIC_API_KEY from the environmentresponse = client.messages.create(    model="claude-opus-5",    max_tokens=200,    output_config={"effort": "low"},    system="You sort courier support messages. Reply with one word: reschedule, "           "address_change, damaged_parcel, complaint, other or unknown.",    messages=[{"role": "user", "content": "please deliver tomorrow after 6, I'm not home"}],)text = "".join(block.text for block in response.content if block.type == "text")print(text)                                    # rescheduleprint(response.stop_reason)                    # end_turnprint(response.usage.input_tokens, response.usage.output_tokens)

The response's content is a list of blocks, because a reply can contain text, thinking and tool calls. You keep only the text blocks. stop_reason tells you why generation stopped, and usage gives the tokens you are billed for.

Here is the same request with OpenAI's Python SDK and its Responses API:

Python
from openai import OpenAIclient = OpenAI()                     # reads OPENAI_API_KEY from the environmentresponse = client.responses.create(    model="gpt-5-mini",    max_output_tokens=200,    reasoning={"effort": "low"},    instructions="You sort courier support messages. Reply with one word: reschedule, "                 "address_change, damaged_parcel, complaint, other or unknown.",    input="please deliver tomorrow after 6, I'm not home",)print(response.output_text)                    # rescheduleprint(response.status)                         # completed (or incomplete)print(response.usage.input_tokens, response.usage.output_tokens)

The ideas are identical; only the names change.

IdeaAnthropic Messages APIOpenAI Responses API
Instructionssysteminstructions
Conversationmessagesinput (a string or a list of turns)
Output limitmax_tokensmax_output_tokens
Thinking controloutput_config={"effort": ...}reasoning={"effort": ...}
Reply texttext blocks in contentoutput_text
Why it stoppedstop_reasonstatus and incomplete_details
Tokens billedusage.input_tokens, usage.output_tokensusage.input_tokens, usage.output_tokens

OpenAI's older Chat Completions API (client.chat.completions.create) is still widely used. It puts the system prompt inside messages with role system, limits output with max_completion_tokens and reports usage.prompt_tokens and usage.completion_tokens. After this lesson, ShipFast hides all of these differences behind one class, so the rest of the code never sees them.

Reading the stop reason

The most ignored field is the most important one for correctness. A reply can come back with status 200 and still be incomplete.

stop_reasonMeaningWhat ShipFast does
end_turnThe model finished on its ownUse the reply
max_tokensIt hit your limit mid-answerTreat as a failure; the JSON is probably cut off
refusalThe model declined for safety reasonsSend the message to a human queue

A reply that stopped at max_tokens often looks almost right: {"intent": "resched. If you only check the status code, you will pass half an answer to your parser. Always check why the model stopped before you trust what it said.

Counting tokens

You pay per token and your limits are in tokens, so you need to count them. Three rules of thumb:

  • English: about 4 characters per token. ShipFast's typical 250-character message is about 60 tokens.
  • Hindi in Devanagari script uses noticeably more tokens for the same meaning, often two to three times as many. A "Hinglish" message in Latin letters sits in between.
  • Numbers, IDs and unusual spellings split into many small tokens. "SF20931847" is several tokens, not one.

For an exact count, ask the provider. Tokenizers differ between providers and even between model generations, so never use one provider's tokenizer library to estimate another's.

Python
count = anthropic.Anthropic().messages.count_tokens(    model="claude-opus-5",    system="You sort courier support messages. ...",    messages=[{"role": "user", "content": "kal shaam 6 baje ke baad deliver karna please"}],)print(count.input_tokens)

Counting is a network call, so ShipFast does not count before every request. It uses a cheap estimate for budget checks (the next two lessons) and reads the exact usage from every response for accounting.

Check your understanding

0 of 3 answered

1.A classification reply arrives with status 200 and stop_reason equal to max_tokens. What should the service do?

2.Why should you not use one provider's tokenizer library to estimate token counts for another provider's model?

3.Your classifier sets max_tokens=20 and high effort. Some replies come back empty with stop_reason of max_tokens. What is the most likely cause?