Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Rate limits and quotas per customer
At 14:05 on a Thursday, a merchant's order system had a bug. For every order update, it sent ShipFast a "customer message", and it retried each one because it did not understand the 202. In ten minutes, ShipFast received 6,000 messages from one business account.
Every one of them was triaged. The provider's rate limit for ShipFast's account, which is shared by all of ShipFast's traffic, was reached within four minutes. From then on, 429 errors hit every customer's messages, the circuit breaker opened, and the fallback model took the load until it too was rate-limited. For about 20 minutes, most real customers' messages were degraded to human triage because one merchant had a bug.
The provider did exactly what it promised. ShipFast had no limits of its own.
The provider's limit is everyone's limit
Providers limit your account, usually in three ways at once: requests per minute, input tokens per minute and output tokens per minute. Exceed any of them and you get 429 until the window resets. The exact numbers depend on your account's tier and the model.
The key word is account. Your 40,000 customers share one limit. Without your own controls, the provider's limit is allocated first-come, first-served, which means the noisiest customer wins. You need three limits of your own, each protecting against a different failure.
Per-customer rate
- Messages per second per customer, with a burst
- Stops floods and loops from one sender
- Rejects with 429 and
Retry-After
Per-customer daily spend
- Model dollars per customer per day
- Stops slow, steady abuse and runaway integrations
- Degrades to human triage instead of rejecting
Global concurrency
- Model calls in flight across the whole service
- Keeps total traffic under the provider limit
- Queues work instead of failing it
Per-customer rate: a token bucket
A token bucket allows short bursts but enforces an average rate. Each customer has a bucket that holds up to burst tokens and refills at rate tokens per second. Each message takes one token. An empty bucket means "too many, try later", and the time until the next token is exactly the Retry-After to send.
1# shipfast/ratelimit.py2import math3import time4from datetime import date56from fastapi import HTTPException789class TokenBucket:10 def __init__(self, rate_per_s: float, burst: int) -> None:11 self.rate, self.capacity = rate_per_s, burst12 self.tokens, self.updated = float(burst), time.monotonic()1314 def take(self) -> float:15 """Take one token. Returns 0.0 if allowed, else seconds until one is free."""16 now = time.monotonic()17 self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)18 self.updated = now19 if self.tokens >= 1:20 self.tokens -= 121 return 0.022 return (1 - self.tokens) / self.rate232425# (messages per second, burst) and daily model spend in USD, per plan26RATES = {"individual": (5 / 60, 10), "business": (1.0, 120)}27DAILY_USD = {"individual": 0.50, "business": 40.00}28_buckets: dict[str, TokenBucket] = {} # in production: Redis, one key per customer293031def check_rate(customer_id: str, plan: str) -> None:32 bucket = _buckets.setdefault(customer_id, TokenBucket(*RATES[plan]))33 wait = bucket.take()34 if wait:35 raise HTTPException(429, "too many messages", headers={"Retry-After": str(math.ceil(wait))})An individual customer may send 10 messages in a burst and then one every 12 seconds, far more than any real person needs. A business account may burst 120 and then send one per second, about 86,000 a day. The merchant with the bug would have been stopped after its first 120 messages, and every later retry would have received a 429 with Retry-After, which well-behaved clients respect.
The in-memory dictionary works for one process. With several instances, each would keep its own buckets and the real limit would multiply by the number of instances. In production, the bucket state lives in Redis, updated atomically with a short Lua script or a tested rate-limit library, and the keys expire after an hour of inactivity so memory does not grow forever.
Per-customer daily spend
Rate limits stop floods, but not a customer who sends long messages steadily all day, or an integration that sends one message a second, every second. A spend quota caps what each customer can cost in model calls per day.
When the quota runs out, ShipFast does not reject the message. A real customer's message still needs an answer. It skips the model and sends the message to human triage, exactly like a provider outage.
1# shipfast/ratelimit.py (continued)2class DailySpend:3 def __init__(self) -> None:4 self._spent: dict[tuple[str, date], float] = {}56 def remaining(self, customer_id: str, plan: str, today: date) -> float:7 return DAILY_USD[plan] - self._spent.get((customer_id, today), 0.0)89 def add(self, customer_id: str, today: date, usd: float) -> None:10 key = (customer_id, today)11 self._spent[key] = self._spent.get(key, 0.0) + usdWiring the three limits in
1# shipfast/api.py (triage part, updated)2import asyncio34from shipfast.ratelimit import DailySpend, check_rate5from shipfast.schemas import Intent, Queue67SPEND = DailySpend()8LLM_SLOTS = asyncio.Semaphore(32) # at most 32 triages talk to the provider at once91011def plan_of(msg: InboundMessage) -> str:12 return "business" if msg.is_business else "individual"131415async def run_triage(msg: InboundMessage, llm, cache) -> None:16 today = today_ist()17 if SPEND.remaining(msg.customer_id, plan_of(msg), today) <= 0:18 RESULTS[msg.message_id] = TriageResult( # over quota: skip the model, keep the message19 message_id=msg.message_id, intent=Intent.UNKNOWN, queue=Queue.HUMAN_TRIAGE,20 priority="normal", degraded=True)21 return22 async with LLM_SLOTS:23 result = await triage(llm, msg, today, cache)24 SPEND.add(msg.customer_id, today, result.cost_usd)25 RESULTS[msg.message_id] = result262728@app.post("/v1/messages", status_code=202)29async def accept_message(msg: InboundMessage, tasks: BackgroundTasks,30 llm=Depends(get_llm), cache=Depends(get_cache)) -> dict:31 check_rate(msg.customer_id, plan_of(msg)) # 429 with Retry-After when over the limit32 if msg.message_id not in RESULTS:33 RESULTS[msg.message_id] = None34 tasks.add_task(run_triage, msg, llm, cache)35 return {"message_id": msg.message_id, "status_url": f"/v1/messages/{msg.message_id}"}The rate check runs first, before any work is recorded, so a flood costs a dictionary lookup per message. The spend check runs in the background job, because the gateway has already got its answer. The semaphore caps the whole process at 32 triages in flight; extra jobs wait for a slot rather than failing. With the job table from Section 4, the worker pool size plays the same role.
The daily limits are generous on purpose. $0.50 is about 40 fully triaged messages, more than any individual sends in a day. $40 is about 3,300 messages for a business. The limits exist to cap accidents, not to ration normal use. Set them from real usage data (ShipFast used the 99.9th percentile of daily spend per customer, times three) and alert when any customer reaches half their quota, so you hear about a runaway integration before it hits the limit.
Check your understanding
0 of 3 answered
1.Why does ShipFast send an over-quota message to human triage instead of rejecting it?
2.ShipFast runs 4 instances, each with the in-memory token bucket. A business account's burst is 120. What is its real burst?
3.Why is the rate check done in the request handler, but the spend check in the background job?