Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you implement rate limiting and request throttling?
What you need to know
Why LLM rate limiting is different
- Unit of cost. A 200-token question and a 50,000-token document summary are both "one request", but the second costs 250× more. Limit tokens, not just requests.
- Duration. A normal API request takes 50 ms; an LLM request takes 2–30 seconds. Concurrency, not request rate, is what fills your GPUs.
The three limits
| Limit | Protects | Typical setting (illustrative) |
|---|---|---|
| Requests per minute (RPM) | Against scripts and abuse | 30 per user |
| Tokens per minute (TPM) | Your budget and provider quota | 20,000 per user |
| Concurrent requests | GPU KV cache and worker slots | 3 per user, 200 per tenant |
Estimating tokens before the call
You do not know the output length in advance. Reserve input_tokens + max_tokens, count input with a tokenizer or the provider's token-counting endpoint, and after the response give back what was not used.
1import time23class TokenBucket:4 def __init__(self, capacity: float, refill_per_sec: float):5 self.capacity = capacity6 self.tokens = capacity7 self.rate = refill_per_sec8 self.last = time.monotonic()910 def try_take(self, amount: float) -> bool:11 now = time.monotonic()12 self.tokens = min(self.capacity, self.tokens + (now - self.last) * self.rate)13 self.last = now14 if self.tokens >= amount:15 self.tokens -= amount16 return True17 return False1819# Per citizen: 20,000 LLM tokens per minute, bursts up to 20,00020tpm = TokenBucket(capacity=20_000, refill_per_sec=20_000 / 60)21estimate = 1_500 + 400 # estimated input + max output tokens22results = [tpm.try_take(estimate) for _ in range(12)]23print(results) # the 11th request is refusedTen requests of 1,900 tokens use 19,000; the eleventh does not fit and is refused until the bucket refills at about 333 tokens a second. In production the same logic runs as an atomic Redis script so all replicas share one bucket.
Admission control and priorities
Behind the limits, keep a bounded queue. If the expected wait is longer than your latency target, reject at once with 429 and a Retry-After header — a fast "try again in 10 seconds" is better than a 60-second spinner. Give interactive traffic a higher priority than batch jobs.
A real-life example
A state government's multilingual chatbot answers questions about a scholarship scheme. On the last day to apply, traffic goes from 50 to 900 requests a second. A few scripts also start scraping it.
The limits work in layers. Per-user RPM (30) stops the scrapers. Per-user TPM (20,000) stops a few users pasting 40-page PDFs from taking the GPU pool. The per-tenant concurrency cap on the self-hosted cluster (1,200 in flight) protects the KV cache. When the queue wait passes 8 seconds, new requests get 429 with Retry-After: 15, and the web app shows "High demand — your question is saved, we will answer in about 15 seconds" and retries once automatically.
Follow-up questions to expect
- "Token bucket or sliding window?" — Token bucket allows controlled bursts and is cheap; a sliding window is stricter but costlier. Token bucket is the usual choice.
- "How do you rate-limit fairly between tenants?" — Per-tenant quotas plus weighted fair queueing, so one big tenant cannot starve small ones.
- "What status code and headers?" —
429 Too Many RequestswithRetry-After, and ideally headers showing remaining quota.