LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you handle rate limits from LLM providers?


Backoff ceilings in seconds, with full jitter0.5124801234attempt 1attempt 3capEach wait is random between zero and the ceiling, and never shorter than Retry-After.
Jitter turns a thousand clients retrying at the same instant into a smooth trickle the provider can absorb.

What you need to know

What providers limit

  • RPM — requests per minute.
  • TPM — tokens per minute, often counted as input plus max_tokens at request time, so a large max_tokens uses quota even if the answer is short.
  • Sometimes concurrency, daily token caps, or separate limits for batch jobs.

Limits are usually per organisation or project and per model, and grow with usage tier. Some providers also sell provisioned throughput — reserved capacity with a guaranteed rate.

Client-side strategy

  • Self-throttle. Keep a token bucket at about 90% of the provider limit. Waiting in your own queue is cheaper than a 429 round trip and keeps the provider's view of you healthy.
  • Right-size max_tokens. If the TPM limit counts max_tokens, setting it to 4,000 when answers are 300 tokens wastes most of your quota.
  • Backoff with full jitter on 429 and 5xx, honouring Retry-After, at most 2–3 attempts, inside a time budget.
  • Separate workloads. Different keys or projects for interactive and batch traffic.
  • Batch APIs. Offline work at about half price, with separate limits.
  • Spread load. Multiple regions, accounts, or a second provider behind the gateway.
  • Prioritise. When close to the limit, interactive requests go first; background work waits.

Backoff delays

With a base of 0.5 s and a cap of 8 s, the maximum waits are 0.5, 1, 2, 4 and 8 seconds. Full jitter picks a random value between 0 and that maximum each time. Without jitter, a thousand clients that failed at the same moment retry at the same moment, and fail again together.

A real-life example

A code assistant for 2,000 engineers runs on a provider limit of 2,000,000 tokens per minute. At 10:15 every morning, after stand-ups, it hits about 300 requests a minute with 7,000 tokens each (5,000 input plus max_tokens of 2,000): 2.1 million TPM — just over the limit. Engineers see "rate limited" errors for ten minutes daily.

The fixes, in order of effect:

  1. max_tokens from 2,000 to 800, since p99 answers are 650 tokens. Counted tokens fall to 5,800 per request: 1.74 million TPM.
  2. Nightly code-review summaries (a batch job that used the same key) move to the batch API with its own quota.
  3. A gateway token bucket at 1.8 million TPM queues the few requests above that for a second or two rather than failing.
  4. A quota increase request is filed when usage reaches 60% of the new headroom, because approval can take days.

Rate-limit errors fall to zero, and nobody had to buy more capacity.

Follow-up questions to expect

  • "Why full jitter rather than fixed backoff?" — Fixed delays keep synchronised clients synchronised; randomness spreads retries out, so the provider sees a smooth load.
  • "What if Retry-After is longer than your timeout?" — Do not wait; fail over to another deployment or tier, or return a clear busy message.
  • "How do you share one quota among many teams?" — Per-team token buckets at the gateway, with priorities and budgets, so no single team can use it all.