LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you handle API rate limits in LangChain?


Three layers between your code and a 429Throttle —InMemoryRateLimiter at 7 requests per secondRetry — only RateLimitErrorand timeouts, with jitterShed load — max_concurrency, cache, fallback modelProvider quota — 500 requests per minute
Retries alone add traffic to a quota that is already full; throttling first is what stops the storm from starting.

What you need to know

Providers limit both requests per minute (RPM) and tokens per minute (TPM). Go over either and you get a 429 "Too Many Requests" response.

Layer 1: throttle before you send

Python
from langchain_core.rate_limiters import InMemoryRateLimiterfrom langchain_openai import ChatOpenAIimport openailimiter = InMemoryRateLimiter(requests_per_second=5, check_every_n_seconds=0.1,                              max_bucket_size=10)llm = ChatOpenAI(model="gpt-5.4-mini", rate_limiter=limiter, max_retries=3)

The limiter counts requests, not tokens, so set it below your provider's request limit with room for the token limit too.

Layer 2: retry what still fails

Python
chain = (prompt | llm | parser).with_retry(    retry_if_exception_type=(openai.RateLimitError, openai.APITimeoutError),    wait_exponential_jitter=True,    stop_after_attempt=3,)

max_retries on the model already retries 429, 5xx and network errors with backoff. with_retry wraps the whole chain. Its default retries on every exception, so limit it to the transient ones; retrying a 400 "bad request" only wastes time.

Layer 3: shed load

  • chain.batch(inputs, config={"max_concurrency": 5}) caps parallel calls.
  • set_llm_cache(InMemoryCache()) (from langchain_core.globals and langchain_core.caches) skips exact repeat calls.
  • llm.with_fallbacks([backup_llm]) switches to another model or provider when the main one keeps failing.

A real-life example

An e-commerce company in Bengaluru re-writes 60,000 product descriptions every night before a sale. Their account allows 500 requests per minute. The first run fired 50 parallel calls, hit 429s within 20 seconds, and the default retries piled on more requests, so the job took 5 hours and 7% of items failed. The fix: an InMemoryRateLimiter at 7 requests per second in the single worker process (about 420 per minute, under the limit), max_concurrency=20, with_retry only for RateLimitError and timeouts, and return_exceptions=True so failures were collected and re-queued. The job finished in about 2.5 hours with no failed items.

Follow-up questions to expect

  • "How do you handle the tokens-per-minute limit?" — Estimate tokens per request, lower concurrency for long inputs, and use the provider's batch API for large offline jobs.
  • "What is jitter and why add it?" — A small random delay added to each backoff, so many clients that failed together do not retry at the same moment.
  • "What if one provider is down for an hour?" — Retries will not help; with_fallbacks to a second provider or a smaller model keeps the service up.