Course Content
FastAPI Essentials
1 sections · 32 lessons
How would you implement rate limiting in FastAPI?
What you need to know
Where to enforce it
| Layer | Good for | Limitation |
|---|---|---|
Edge (nginx limit_req, API gateway, Cloudflare) | Per-IP floods, bots, cheap rejection | Does not know your plans or tenants |
| FastAPI dependency + Redis | Per-key, per-plan business limits | Every request touches Redis (about 1 ms) |
In-process counter or slowapi with memory storage | A single-process demo | Wrong as soon as you run 2+ workers |
The last row is the classic trap. With 4 workers, each keeps its own count, so a "60 per minute" limit really allows up to 240. The counter must live in shared storage.
A per-plan limiter as a dependency
1import time2from typing import Annotated3from fastapi import Depends, FastAPI, HTTPException4from redis.asyncio import Redis56redis = Redis.from_url("redis://localhost:6379")7app = FastAPI()8LIMITS = {"free": 5, "pro": 60} # requests per minute910async def rate_limit(key: Annotated[tuple[str, str], Depends(api_key)]) -> str:11 tenant, plan = key # api_key dependency looks up the tenant12 bucket = f"rl:{tenant}:{int(time.time() // 60)}"13 used = await redis.incr(bucket) # atomic across all workers and pods14 if used == 1:15 await redis.expire(bucket, 120)16 if used > LIMITS[plan]:17 retry = 60 - int(time.time() % 60)18 raise HTTPException(429, "Rate limit exceeded", headers={"Retry-After": str(retry)})19 return tenant2021@app.post("/v1/chat")22async def chat(prompt: str, tenant: Annotated[str, Depends(rate_limit)]):23 return {"tenant": tenant, "answer": "..."}Run with an in-memory Redis stand-in (fakeredis), real results:
free plan, 7 calls: [200, 200, 200, 200, 200, 429, 429]429 {'detail': 'Rate limit exceeded'} Retry-After: 5pro plan: 200This is a fixed window: one counter per tenant per minute. Redis INCR is atomic, so two workers can never both read "59" and both allow request 60. Fixed windows allow a burst at the boundary (5 calls at 12:00:59 and 5 more at 12:01:00); a sliding window or token bucket smooths this, and libraries such as slowapi or fastapi-limiter with Redis storage implement them for you.
LLM-specific limits
- Tokens per minute. After each call, add the real token usage:
await redis.incrby(f"tok:{tenant}:{minute}", usage.total_tokens). Check the budget before the next call. - Concurrency. An
asyncio.Semaphore(20)per worker, or a Redis counter per tenant, caps how many generations run at once. This protects the GPU or your provider quota. - Cost caps. A daily spend limit per tenant stops a runaway script from running up a large bill overnight.
A real-life example
A startup sells an LLM writing API with free and pro plans. A free-tier user's script sent 50 requests per second, each with a 6,000-token prompt. The team's first limiter used slowapi with its default in-memory storage, and they ran 6 workers across 2 pods: the real limit was 12 times the configured one, and the user burned through the month's provider budget for that plan in two hours.
They moved the counter to Redis, added a tokens-per-minute budget (free: 20,000; pro: 400,000) and a limit of 3 concurrent generations per free key. Nginx limit_req handles IP floods before Python. Well-behaved clients read Retry-After and back off, so the 429 rate fell quickly after the SDK was updated to respect it.
Follow-up questions to expect
- "Why per API key and not per IP?" — Many users share one IP (an office, a mobile carrier's NAT), and one attacker can use many IPs. The key is who pays and who has a plan.
- "What if Redis is down?" — Decide in advance: fail open (allow requests, risk cost) or fail closed (reject, risk outage). Most teams fail open with a local fallback limit and an alert.
- "What headers help clients?" —
Retry-Afteron 429, and optionallyRateLimit-LimitandRateLimit-Remainingstyle headers on every response.