Building AI Features in Python Backends

Rate limits and quotas per customer


At 14:05 on a Thursday, a merchant's order system had a bug. For every order update, it sent ShipFast a "customer message", and it retried each one because it did not understand the 202. In ten minutes, ShipFast received 6,000 messages from one business account.

Every one of them was triaged. The provider's rate limit for ShipFast's account, which is shared by all of ShipFast's traffic, was reached within four minutes. From then on, 429 errors hit every customer's messages, the circuit breaker opened, and the fallback model took the load until it too was rate-limited. For about 20 minutes, most real customers' messages were degraded to human triage because one merchant had a bug.

The provider did exactly what it promised. ShipFast had no limits of its own.

Three limits in front of one shared provider limitTokenbucket percustomer: 429Daily spend quota:degrade, not rejectSemaphore:32 triagesin flightProvider limit,shared by everyoneA merchant's 6,000-message loop stopped after its burst of 120.
The provider limits your account, not your customers, so without your own limits the noisiest customer takes everyone's capacity.

The provider's limit is everyone's limit

Providers limit your account, usually in three ways at once: requests per minute, input tokens per minute and output tokens per minute. Exceed any of them and you get 429 until the window resets. The exact numbers depend on your account's tier and the model.

The key word is account. Your 40,000 customers share one limit. Without your own controls, the provider's limit is allocated first-come, first-served, which means the noisiest customer wins. You need three limits of your own, each protecting against a different failure.

Per-customer rate

  • Messages per second per customer, with a burst
  • Stops floods and loops from one sender
  • Rejects with 429 and Retry-After

Per-customer daily spend

  • Model dollars per customer per day
  • Stops slow, steady abuse and runaway integrations
  • Degrades to human triage instead of rejecting

Global concurrency

  • Model calls in flight across the whole service
  • Keeps total traffic under the provider limit
  • Queues work instead of failing it

Per-customer rate: a token bucket

A token bucket allows short bursts but enforces an average rate. Each customer has a bucket that holds up to burst tokens and refills at rate tokens per second. Each message takes one token. An empty bucket means "too many, try later", and the time until the next token is exactly the Retry-After to send.

Python
# shipfast/ratelimit.pyimport mathimport timefrom datetime import datefrom fastapi import HTTPExceptionclass TokenBucket:    def __init__(self, rate_per_s: float, burst: int) -> None:        self.rate, self.capacity = rate_per_s, burst        self.tokens, self.updated = float(burst), time.monotonic()    def take(self) -> float:        """Take one token. Returns 0.0 if allowed, else seconds until one is free."""        now = time.monotonic()        self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)        self.updated = now        if self.tokens >= 1:            self.tokens -= 1            return 0.0        return (1 - self.tokens) / self.rate# (messages per second, burst) and daily model spend in USD, per planRATES = {"individual": (5 / 60, 10), "business": (1.0, 120)}DAILY_USD = {"individual": 0.50, "business": 40.00}_buckets: dict[str, TokenBucket] = {}      # in production: Redis, one key per customerdef check_rate(customer_id: str, plan: str) -> None:    bucket = _buckets.setdefault(customer_id, TokenBucket(*RATES[plan]))    wait = bucket.take()    if wait:        raise HTTPException(429, "too many messages", headers={"Retry-After": str(math.ceil(wait))})

An individual customer may send 10 messages in a burst and then one every 12 seconds, far more than any real person needs. A business account may burst 120 and then send one per second, about 86,000 a day. The merchant with the bug would have been stopped after its first 120 messages, and every later retry would have received a 429 with Retry-After, which well-behaved clients respect.

The in-memory dictionary works for one process. With several instances, each would keep its own buckets and the real limit would multiply by the number of instances. In production, the bucket state lives in Redis, updated atomically with a short Lua script or a tested rate-limit library, and the keys expire after an hour of inactivity so memory does not grow forever.

Per-customer daily spend

Rate limits stop floods, but not a customer who sends long messages steadily all day, or an integration that sends one message a second, every second. A spend quota caps what each customer can cost in model calls per day.

When the quota runs out, ShipFast does not reject the message. A real customer's message still needs an answer. It skips the model and sends the message to human triage, exactly like a provider outage.

Python
# shipfast/ratelimit.py (continued)class DailySpend:    def __init__(self) -> None:        self._spent: dict[tuple[str, date], float] = {}    def remaining(self, customer_id: str, plan: str, today: date) -> float:        return DAILY_USD[plan] - self._spent.get((customer_id, today), 0.0)    def add(self, customer_id: str, today: date, usd: float) -> None:        key = (customer_id, today)        self._spent[key] = self._spent.get(key, 0.0) + usd

Wiring the three limits in

Python
# shipfast/api.py (triage part, updated)import asynciofrom shipfast.ratelimit import DailySpend, check_ratefrom shipfast.schemas import Intent, QueueSPEND = DailySpend()LLM_SLOTS = asyncio.Semaphore(32)        # at most 32 triages talk to the provider at oncedef plan_of(msg: InboundMessage) -> str:    return "business" if msg.is_business else "individual"async def run_triage(msg: InboundMessage, llm, cache) -> None:    today = today_ist()    if SPEND.remaining(msg.customer_id, plan_of(msg), today) <= 0:        RESULTS[msg.message_id] = TriageResult(      # over quota: skip the model, keep the message            message_id=msg.message_id, intent=Intent.UNKNOWN, queue=Queue.HUMAN_TRIAGE,            priority="normal", degraded=True)        return    async with LLM_SLOTS:        result = await triage(llm, msg, today, cache)    SPEND.add(msg.customer_id, today, result.cost_usd)    RESULTS[msg.message_id] = result@app.post("/v1/messages", status_code=202)async def accept_message(msg: InboundMessage, tasks: BackgroundTasks,                         llm=Depends(get_llm), cache=Depends(get_cache)) -> dict:    check_rate(msg.customer_id, plan_of(msg))    # 429 with Retry-After when over the limit    if msg.message_id not in RESULTS:        RESULTS[msg.message_id] = None        tasks.add_task(run_triage, msg, llm, cache)    return {"message_id": msg.message_id, "status_url": f"/v1/messages/{msg.message_id}"}

The rate check runs first, before any work is recorded, so a flood costs a dictionary lookup per message. The spend check runs in the background job, because the gateway has already got its answer. The semaphore caps the whole process at 32 triages in flight; extra jobs wait for a slot rather than failing. With the job table from Section 4, the worker pool size plays the same role.

The daily limits are generous on purpose. $0.50 is about 40 fully triaged messages, more than any individual sends in a day. $40 is about 3,300 messages for a business. The limits exist to cap accidents, not to ration normal use. Set them from real usage data (ShipFast used the 99.9th percentile of daily spend per customer, times three) and alert when any customer reaches half their quota, so you hear about a runaway integration before it hits the limit.

Check your understanding

0 of 3 answered

1.Why does ShipFast send an over-quota message to human triage instead of rejecting it?

2.ShipFast runs 4 instances, each with the in-memory token bucket. A business account's burst is 120. What is its real burst?

3.Why is the rate check done in the request handler, but the spend check in the background job?