Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Anatomy of a request: messages, parameters, tokens
In the last section you saw that a model call is an HTTPS request. Now you need to know that request well enough to control it. Every field changes something you care about: what the model sees, how long it can write, how long it takes and what you pay.
ShipFast's first classifier call was copied from a blog post. It had no limit on output length, sent the instructions as part of the user's text, ignored the response's stop reason and never looked at the token counts. It worked, but nobody could say what it cost or why it sometimes returned half a sentence. This lesson takes a request apart so that does not happen to you.
The parts of a request
A request has three kinds of content and a handful of settings.
- System prompt — your instructions: who the model is, what the task is, the allowed labels, the rules. This is written by you and is the same for every call of one feature.
- Messages — the conversation, as a list of turns with a
roleofuserorassistant. For ShipFast, this is usually one user turn holding the customer's message. In the repair step of Section 3, you add the model's previous reply as anassistantturn and your correction as a newuserturn. - Model — which model runs the request. Keep it in configuration, never hard-coded in the call, because you will change it.
- max_tokens — the most tokens the model may generate. It is a hard stop, not a target. It is also your cost ceiling for output.
- Effort or reasoning settings — many current models can think before answering. More thinking can raise accuracy on hard tasks, but thinking tokens are generated output: they count toward
max_tokens, add latency and are billed. For short tasks like ShipFast's classifier, low effort is the right starting point.
Some providers and older models also accept temperature, which controls how random the sampling is. Several current models no longer accept it. Do not build anything that depends on it.
The same call with two SDKs
Here is ShipFast's classifier request with Anthropic's Python SDK. It is synchronous to keep it short; ShipFast itself uses the async client.
1import anthropic23client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment45response = client.messages.create(6 model="claude-opus-5",7 max_tokens=200,8 output_config={"effort": "low"},9 system="You sort courier support messages. Reply with one word: reschedule, "10 "address_change, damaged_parcel, complaint, other or unknown.",11 messages=[{"role": "user", "content": "please deliver tomorrow after 6, I'm not home"}],12)13text = "".join(block.text for block in response.content if block.type == "text")14print(text) # reschedule15print(response.stop_reason) # end_turn16print(response.usage.input_tokens, response.usage.output_tokens)The response's content is a list of blocks, because a reply can contain text, thinking and tool calls. You keep only the text blocks. stop_reason tells you why generation stopped, and usage gives the tokens you are billed for.
Here is the same request with OpenAI's Python SDK and its Responses API:
1from openai import OpenAI23client = OpenAI() # reads OPENAI_API_KEY from the environment45response = client.responses.create(6 model="gpt-5-mini",7 max_output_tokens=200,8 reasoning={"effort": "low"},9 instructions="You sort courier support messages. Reply with one word: reschedule, "10 "address_change, damaged_parcel, complaint, other or unknown.",11 input="please deliver tomorrow after 6, I'm not home",12)13print(response.output_text) # reschedule14print(response.status) # completed (or incomplete)15print(response.usage.input_tokens, response.usage.output_tokens)The ideas are identical; only the names change.
| Idea | Anthropic Messages API | OpenAI Responses API |
|---|---|---|
| Instructions | system | instructions |
| Conversation | messages | input (a string or a list of turns) |
| Output limit | max_tokens | max_output_tokens |
| Thinking control | output_config={"effort": ...} | reasoning={"effort": ...} |
| Reply text | text blocks in content | output_text |
| Why it stopped | stop_reason | status and incomplete_details |
| Tokens billed | usage.input_tokens, usage.output_tokens | usage.input_tokens, usage.output_tokens |
OpenAI's older Chat Completions API (client.chat.completions.create) is still widely used. It puts the system prompt inside messages with role system, limits output with max_completion_tokens and reports usage.prompt_tokens and usage.completion_tokens. After this lesson, ShipFast hides all of these differences behind one class, so the rest of the code never sees them.
Reading the stop reason
The most ignored field is the most important one for correctness. A reply can come back with status 200 and still be incomplete.
stop_reason | Meaning | What ShipFast does |
|---|---|---|
end_turn | The model finished on its own | Use the reply |
max_tokens | It hit your limit mid-answer | Treat as a failure; the JSON is probably cut off |
refusal | The model declined for safety reasons | Send the message to a human queue |
A reply that stopped at max_tokens often looks almost right: {"intent": "resched. If you only check the status code, you will pass half an answer to your parser. Always check why the model stopped before you trust what it said.
Counting tokens
You pay per token and your limits are in tokens, so you need to count them. Three rules of thumb:
- English: about 4 characters per token. ShipFast's typical 250-character message is about 60 tokens.
- Hindi in Devanagari script uses noticeably more tokens for the same meaning, often two to three times as many. A "Hinglish" message in Latin letters sits in between.
- Numbers, IDs and unusual spellings split into many small tokens. "SF20931847" is several tokens, not one.
For an exact count, ask the provider. Tokenizers differ between providers and even between model generations, so never use one provider's tokenizer library to estimate another's.
1count = anthropic.Anthropic().messages.count_tokens(2 model="claude-opus-5",3 system="You sort courier support messages. ...",4 messages=[{"role": "user", "content": "kal shaam 6 baje ke baad deliver karna please"}],5)6print(count.input_tokens)Counting is a network call, so ShipFast does not count before every request. It uses a cheap estimate for budget checks (the next two lessons) and reads the exact usage from every response for accounting.
Check your understanding
0 of 3 answered
1.A classification reply arrives with status 200 and stop_reason equal to max_tokens. What should the service do?
2.Why should you not use one provider's tokenizer library to estimate token counts for another provider's model?
3.Your classifier sets max_tokens=20 and high effort. Some replies come back empty with stop_reason of max_tokens. What is the most likely cause?