Building AI Features in Python Backends

Token limits, cost accounting and budgets per request


A model call with no limits is an open invoice. ShipFast learned this in its second week. A merchant's system forwarded a whole order export, 38,000 characters, as a "customer message". The classifier read all of it. The extractor read it again. The repair step read it a third time with the failed reply attached. One message cost $0.41, about a hundred times the average, and there was nothing in the code to stop it.

Nothing about that bug is specific to AI. It is an unbounded input driving an unbounded cost. The fix is the same as for any expensive dependency: limit what goes in, limit what comes out, and keep a running total per request that you check before each call.

One long message against a 0.03 USD budget0.0060.0130.0210.0320123classifyextractrepairdraft,worst caseRunning total in USD. The draft's worst case would cross 0.03, so it never starts.
The budget is checked with the worst case before each call, so the call that would cross the line is never made.

Three limits, three different jobs

Context window

  • The provider's maximum: input plus output
  • Often 200,000 to 1,000,000 tokens today
  • Not a target; a ceiling you should never approach in a request path

Your input limit

  • Set by you, before any model call
  • ShipFast: 2,000 characters per message
  • Rejects bad input cheaply, with a 422

max_tokens

  • Set by you, per call
  • Caps output length, latency and output cost
  • Classify: 200. Extract: 300. Draft: 250

The context window is large, and that makes it dangerous. A model that accepts 200,000 tokens will happily bill you for 200,000 tokens. Your own input limit is the real defence. In ShipFast, the request model enforces it, so an oversized message never reaches a model at all.

Python
from pydantic import BaseModel, Fieldclass InboundMessage(BaseModel):    message_id: str = Field(min_length=1, max_length=64)    customer_id: str    is_business: bool = False    text: str = Field(min_length=1, max_length=2_000)

FastAPI returns a 422 for the 38,000-character export before any code of yours runs. Real customers almost never write more than 2,000 characters; the rare long email is sent to a human queue by the gateway instead.

Pricing every call

Cost is input tokens times the input price, plus output tokens times the output price. Prices are quoted per million tokens and differ a lot between models, so ShipFast keeps them in one table.

Python
# shipfast/config.pyimport osfrom dataclasses import dataclass# USD per million tokens: (input, output). Check the provider's pricing page.PRICES_PER_MTOK: dict[str, tuple[float, float]] = {    "claude-opus-5": (5.00, 25.00),    "claude-sonnet-5": (2.00, 10.00),    "claude-haiku-4-5": (1.00, 5.00),}@dataclass(frozen=True)class Settings:    llm_model: str = os.getenv("SHIPFAST_LLM_MODEL", "claude-opus-5")    llm_fallback_model: str = os.getenv("SHIPFAST_LLM_FALLBACK_MODEL", "claude-sonnet-5")    llm_timeout_s: float = float(os.getenv("SHIPFAST_LLM_TIMEOUT_S", "10"))    max_message_chars: int = 2_000    request_budget_usd: float = 0.03settings = Settings()

These are list prices at the time of writing; prices change, so update the table from the provider's pricing page when you upgrade. Because LLMResult.cost_usd looks up the table by model name, an unknown model raises a KeyError on the first call. That is on purpose: you cannot deploy a model you have not priced.

Now put numbers on ShipFast's three calls. The prompts and outputs below are the averages from the team's logs.

CallInput tokensOutput tokensOpus 5Sonnet 5Haiku 4.5
Classify41338$0.0030$0.0012$0.0006
Extract47085$0.0045$0.0018$0.0009
Draft43095$0.0045$0.0018$0.0009
Reschedule message (all three)$0.0120$0.0048$0.0024

The cheapest model is five times cheaper per message. Whether it is good enough is a different question, and only an evaluation set can answer it. ShipFast's evaluations in Section 5 showed the smallest model was as accurate as the largest on classification but made more mistakes with relative dates in extraction. So the sensible plan is per feature, not one model for everything. Compare models on cost per correct result, not on price per token.

A budget per request

A limit on each call is not enough, because one request can make several calls: classify, extract, a repair, a draft. ShipFast gives each incoming message a budget and checks it before every call.

Python
# shipfast/budget.pyfrom dataclasses import dataclass, fieldfrom shipfast.config import PRICES_PER_MTOKfrom shipfast.llm import LLMResultclass BudgetExceeded(Exception):    passdef estimate_tokens(text: str) -> int:    # About 4 characters per token for English; Hindi and emoji use far more.    # Dividing by 3 is deliberately pessimistic. Use count_tokens when it matters.    return len(text) // 3 + 1@dataclassclass RequestBudget:    max_usd: float    spent_usd: float = 0.0    calls: list[LLMResult] = field(default_factory=list)    def check(self, model: str, prompt_text: str, max_tokens: int) -> None:        price_in, price_out = PRICES_PER_MTOK[model]        worst_case = (estimate_tokens(prompt_text) * price_in + max_tokens * price_out) / 1_000_000        if self.spent_usd + worst_case > self.max_usd:            raise BudgetExceeded(                f"spent {self.spent_usd:.4f} + worst case {worst_case:.4f} > {self.max_usd}")    def charge(self, result: LLMResult) -> None:        self.calls.append(result)        self.spent_usd += result.cost_usd

check() runs before a call and uses the worst case: a pessimistic token estimate for the input, and the full max_tokens for the output. charge() runs after, with the real usage numbers. The difference is important. Checking with actual cost after the call would let one expensive call through; checking the worst case first means no single call can push a request over its budget.

ShipFast's budget is $0.03 per message. A normal reschedule message spends $0.012 on Opus 5. A message that needs one repair on extraction spends about $0.018. The pathological export would have been blocked long before $0.41. When the budget runs out, the pipeline raises BudgetExceeded and the message goes to a human queue, which Section 4 wires up.

Watching the total

Per-request budgets stop one bad request. You also need to see the trend. Because every call logs feature, model and cost_usd, a daily query over your logs gives cost per feature per day. ShipFast adds two alerts: daily spend over 130% of the 7-day average, and any single request over $0.025. The first catches a prompt change that doubled token use; the second catches a new kind of input nobody anticipated.

Check your understanding

0 of 3 answered

1.Why does RequestBudget.check() use max_tokens for the output instead of the expected output length?

2.A model is five times cheaper per token than another. When is it not the cheaper choice for ShipFast?

3.A merchant sends 38,000 characters as a customer message. Where is the cheapest place to stop it?