LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you implement rate limiting and request throttling?


What you need to know

Why LLM rate limiting is different

  • Unit of cost. A 200-token question and a 50,000-token document summary are both "one request", but the second costs 250× more. Limit tokens, not just requests.
  • Duration. A normal API request takes 50 ms; an LLM request takes 2–30 seconds. Concurrency, not request rate, is what fills your GPUs.

The three limits

LimitProtectsTypical setting (illustrative)
Requests per minute (RPM)Against scripts and abuse30 per user
Tokens per minute (TPM)Your budget and provider quota20,000 per user
Concurrent requestsGPU KV cache and worker slots3 per user, 200 per tenant

Estimating tokens before the call

You do not know the output length in advance. Reserve input_tokens + max_tokens, count input with a tokenizer or the provider's token-counting endpoint, and after the response give back what was not used.

Python
import timeclass TokenBucket:    def __init__(self, capacity: float, refill_per_sec: float):        self.capacity = capacity        self.tokens = capacity        self.rate = refill_per_sec        self.last = time.monotonic()    def try_take(self, amount: float) -> bool:        now = time.monotonic()        self.tokens = min(self.capacity, self.tokens + (now - self.last) * self.rate)        self.last = now        if self.tokens >= amount:            self.tokens -= amount            return True        return False# Per citizen: 20,000 LLM tokens per minute, bursts up to 20,000tpm = TokenBucket(capacity=20_000, refill_per_sec=20_000 / 60)estimate = 1_500 + 400            # estimated input + max output tokensresults = [tpm.try_take(estimate) for _ in range(12)]print(results)                    # the 11th request is refused

Ten requests of 1,900 tokens use 19,000; the eleventh does not fit and is refused until the bucket refills at about 333 tokens a second. In production the same logic runs as an atomic Redis script so all replicas share one bucket.

Admission control and priorities

Behind the limits, keep a bounded queue. If the expected wait is longer than your latency target, reject at once with 429 and a Retry-After header — a fast "try again in 10 seconds" is better than a 60-second spinner. Give interactive traffic a higher priority than batch jobs.

A real-life example

A state government's multilingual chatbot answers questions about a scholarship scheme. On the last day to apply, traffic goes from 50 to 900 requests a second. A few scripts also start scraping it.

The limits work in layers. Per-user RPM (30) stops the scrapers. Per-user TPM (20,000) stops a few users pasting 40-page PDFs from taking the GPU pool. The per-tenant concurrency cap on the self-hosted cluster (1,200 in flight) protects the KV cache. When the queue wait passes 8 seconds, new requests get 429 with Retry-After: 15, and the web app shows "High demand — your question is saved, we will answer in about 15 seconds" and retries once automatically.

Follow-up questions to expect

  • "Token bucket or sliding window?" — Token bucket allows controlled bursts and is cheap; a sliding window is stricter but costlier. Token bucket is the usual choice.
  • "How do you rate-limit fairly between tenants?" — Per-tenant quotas plus weighted fair queueing, so one big tenant cannot starve small ones.
  • "What status code and headers?" — 429 Too Many Requests with Retry-After, and ideally headers showing remaining quota.