FastAPI Essentials

Course Content

FastAPI Essentials

1 sections · 32 lessons

How would you implement rate limiting in FastAPI?


Free plan, five calls a minute, one shared counter123456701234565thcall: 2006thcall: 429Redis INCR on rl:tenant:minute is atomic across every worker and pod.
A counter inside each worker would have allowed five calls per worker; the limit is only real when the count is shared.

What you need to know

Where to enforce it

LayerGood forLimitation
Edge (nginx limit_req, API gateway, Cloudflare)Per-IP floods, bots, cheap rejectionDoes not know your plans or tenants
FastAPI dependency + RedisPer-key, per-plan business limitsEvery request touches Redis (about 1 ms)
In-process counter or slowapi with memory storageA single-process demoWrong as soon as you run 2+ workers

The last row is the classic trap. With 4 workers, each keeps its own count, so a "60 per minute" limit really allows up to 240. The counter must live in shared storage.

A per-plan limiter as a dependency

Python
import timefrom typing import Annotatedfrom fastapi import Depends, FastAPI, HTTPExceptionfrom redis.asyncio import Redisredis = Redis.from_url("redis://localhost:6379")app = FastAPI()LIMITS = {"free": 5, "pro": 60}                    # requests per minuteasync def rate_limit(key: Annotated[tuple[str, str], Depends(api_key)]) -> str:    tenant, plan = key                              # api_key dependency looks up the tenant    bucket = f"rl:{tenant}:{int(time.time() // 60)}"    used = await redis.incr(bucket)                 # atomic across all workers and pods    if used == 1:        await redis.expire(bucket, 120)    if used > LIMITS[plan]:        retry = 60 - int(time.time() % 60)        raise HTTPException(429, "Rate limit exceeded", headers={"Retry-After": str(retry)})    return tenant@app.post("/v1/chat")async def chat(prompt: str, tenant: Annotated[str, Depends(rate_limit)]):    return {"tenant": tenant, "answer": "..."}

Run with an in-memory Redis stand-in (fakeredis), real results:

Text
free plan, 7 calls: [200, 200, 200, 200, 200, 429, 429]429 {'detail': 'Rate limit exceeded'} Retry-After: 5pro plan: 200

This is a fixed window: one counter per tenant per minute. Redis INCR is atomic, so two workers can never both read "59" and both allow request 60. Fixed windows allow a burst at the boundary (5 calls at 12:00:59 and 5 more at 12:01:00); a sliding window or token bucket smooths this, and libraries such as slowapi or fastapi-limiter with Redis storage implement them for you.

LLM-specific limits

  • Tokens per minute. After each call, add the real token usage: await redis.incrby(f"tok:{tenant}:{minute}", usage.total_tokens). Check the budget before the next call.
  • Concurrency. An asyncio.Semaphore(20) per worker, or a Redis counter per tenant, caps how many generations run at once. This protects the GPU or your provider quota.
  • Cost caps. A daily spend limit per tenant stops a runaway script from running up a large bill overnight.

A real-life example

A startup sells an LLM writing API with free and pro plans. A free-tier user's script sent 50 requests per second, each with a 6,000-token prompt. The team's first limiter used slowapi with its default in-memory storage, and they ran 6 workers across 2 pods: the real limit was 12 times the configured one, and the user burned through the month's provider budget for that plan in two hours.

They moved the counter to Redis, added a tokens-per-minute budget (free: 20,000; pro: 400,000) and a limit of 3 concurrent generations per free key. Nginx limit_req handles IP floods before Python. Well-behaved clients read Retry-After and back off, so the 429 rate fell quickly after the SDK was updated to respect it.

Follow-up questions to expect

  • "Why per API key and not per IP?" — Many users share one IP (an office, a mobile carrier's NAT), and one attacker can use many IPs. The key is who pays and who has a plan.
  • "What if Redis is down?" — Decide in advance: fail open (allow requests, risk cost) or fail closed (reject, risk outage). Most teams fail open with a local fallback limit and an alert.
  • "What headers help clients?" — Retry-After on 429, and optionally RateLimit-Limit and RateLimit-Remaining style headers on every response.