Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Token limits, cost accounting and budgets per request
A model call with no limits is an open invoice. ShipFast learned this in its second week. A merchant's system forwarded a whole order export, 38,000 characters, as a "customer message". The classifier read all of it. The extractor read it again. The repair step read it a third time with the failed reply attached. One message cost $0.41, about a hundred times the average, and there was nothing in the code to stop it.
Nothing about that bug is specific to AI. It is an unbounded input driving an unbounded cost. The fix is the same as for any expensive dependency: limit what goes in, limit what comes out, and keep a running total per request that you check before each call.
Three limits, three different jobs
Context window
- The provider's maximum: input plus output
- Often 200,000 to 1,000,000 tokens today
- Not a target; a ceiling you should never approach in a request path
Your input limit
- Set by you, before any model call
- ShipFast: 2,000 characters per message
- Rejects bad input cheaply, with a 422
max_tokens
- Set by you, per call
- Caps output length, latency and output cost
- Classify: 200. Extract: 300. Draft: 250
The context window is large, and that makes it dangerous. A model that accepts 200,000 tokens will happily bill you for 200,000 tokens. Your own input limit is the real defence. In ShipFast, the request model enforces it, so an oversized message never reaches a model at all.
1from pydantic import BaseModel, Field234class InboundMessage(BaseModel):5 message_id: str = Field(min_length=1, max_length=64)6 customer_id: str7 is_business: bool = False8 text: str = Field(min_length=1, max_length=2_000)FastAPI returns a 422 for the 38,000-character export before any code of yours runs. Real customers almost never write more than 2,000 characters; the rare long email is sent to a human queue by the gateway instead.
Pricing every call
Cost is input tokens times the input price, plus output tokens times the output price. Prices are quoted per million tokens and differ a lot between models, so ShipFast keeps them in one table.
1# shipfast/config.py2import os3from dataclasses import dataclass45# USD per million tokens: (input, output). Check the provider's pricing page.6PRICES_PER_MTOK: dict[str, tuple[float, float]] = {7 "claude-opus-5": (5.00, 25.00),8 "claude-sonnet-5": (2.00, 10.00),9 "claude-haiku-4-5": (1.00, 5.00),10}111213@dataclass(frozen=True)14class Settings:15 llm_model: str = os.getenv("SHIPFAST_LLM_MODEL", "claude-opus-5")16 llm_fallback_model: str = os.getenv("SHIPFAST_LLM_FALLBACK_MODEL", "claude-sonnet-5")17 llm_timeout_s: float = float(os.getenv("SHIPFAST_LLM_TIMEOUT_S", "10"))18 max_message_chars: int = 2_00019 request_budget_usd: float = 0.03202122settings = Settings()These are list prices at the time of writing; prices change, so update the table from the provider's pricing page when you upgrade. Because LLMResult.cost_usd looks up the table by model name, an unknown model raises a KeyError on the first call. That is on purpose: you cannot deploy a model you have not priced.
Now put numbers on ShipFast's three calls. The prompts and outputs below are the averages from the team's logs.
| Call | Input tokens | Output tokens | Opus 5 | Sonnet 5 | Haiku 4.5 |
|---|---|---|---|---|---|
| Classify | 413 | 38 | $0.0030 | $0.0012 | $0.0006 |
| Extract | 470 | 85 | $0.0045 | $0.0018 | $0.0009 |
| Draft | 430 | 95 | $0.0045 | $0.0018 | $0.0009 |
| Reschedule message (all three) | $0.0120 | $0.0048 | $0.0024 |
The cheapest model is five times cheaper per message. Whether it is good enough is a different question, and only an evaluation set can answer it. ShipFast's evaluations in Section 5 showed the smallest model was as accurate as the largest on classification but made more mistakes with relative dates in extraction. So the sensible plan is per feature, not one model for everything. Compare models on cost per correct result, not on price per token.
A budget per request
A limit on each call is not enough, because one request can make several calls: classify, extract, a repair, a draft. ShipFast gives each incoming message a budget and checks it before every call.
1# shipfast/budget.py2from dataclasses import dataclass, field34from shipfast.config import PRICES_PER_MTOK5from shipfast.llm import LLMResult678class BudgetExceeded(Exception):9 pass101112def estimate_tokens(text: str) -> int:13 # About 4 characters per token for English; Hindi and emoji use far more.14 # Dividing by 3 is deliberately pessimistic. Use count_tokens when it matters.15 return len(text) // 3 + 1161718@dataclass19class RequestBudget:20 max_usd: float21 spent_usd: float = 0.022 calls: list[LLMResult] = field(default_factory=list)2324 def check(self, model: str, prompt_text: str, max_tokens: int) -> None:25 price_in, price_out = PRICES_PER_MTOK[model]26 worst_case = (estimate_tokens(prompt_text) * price_in + max_tokens * price_out) / 1_000_00027 if self.spent_usd + worst_case > self.max_usd:28 raise BudgetExceeded(29 f"spent {self.spent_usd:.4f} + worst case {worst_case:.4f} > {self.max_usd}")3031 def charge(self, result: LLMResult) -> None:32 self.calls.append(result)33 self.spent_usd += result.cost_usdcheck() runs before a call and uses the worst case: a pessimistic token estimate for the input, and the full max_tokens for the output. charge() runs after, with the real usage numbers. The difference is important. Checking with actual cost after the call would let one expensive call through; checking the worst case first means no single call can push a request over its budget.
ShipFast's budget is $0.03 per message. A normal reschedule message spends $0.012 on Opus 5. A message that needs one repair on extraction spends about $0.018. The pathological export would have been blocked long before $0.41. When the budget runs out, the pipeline raises BudgetExceeded and the message goes to a human queue, which Section 4 wires up.
Watching the total
Per-request budgets stop one bad request. You also need to see the trend. Because every call logs feature, model and cost_usd, a daily query over your logs gives cost per feature per day. ShipFast adds two alerts: daily spend over 130% of the 7-day average, and any single request over $0.025. The first catches a prompt change that doubled token use; the second catches a new kind of input nobody anticipated.
Check your understanding
0 of 3 answered
1.Why does RequestBudget.check() use max_tokens for the output instead of the expected output length?
2.A model is five times cheaper per token than another. When is it not the cheaper choice for ShipFast?
3.A merchant sends 38,000 characters as a customer message. Where is the cheapest place to stop it?