Course Content
LangChain Mastery
7 sections · 109 lessons
How do you handle API rate limits in LangChain?
What you need to know
Providers limit both requests per minute (RPM) and tokens per minute (TPM). Go over either and you get a 429 "Too Many Requests" response.
Layer 1: throttle before you send
1from langchain_core.rate_limiters import InMemoryRateLimiter2from langchain_openai import ChatOpenAI3import openai45limiter = InMemoryRateLimiter(requests_per_second=5, check_every_n_seconds=0.1,6 max_bucket_size=10)7llm = ChatOpenAI(model="gpt-5.4-mini", rate_limiter=limiter, max_retries=3)The limiter counts requests, not tokens, so set it below your provider's request limit with room for the token limit too.
Layer 2: retry what still fails
1chain = (prompt | llm | parser).with_retry(2 retry_if_exception_type=(openai.RateLimitError, openai.APITimeoutError),3 wait_exponential_jitter=True,4 stop_after_attempt=3,5)max_retries on the model already retries 429, 5xx and network errors with backoff. with_retry wraps the whole chain. Its default retries on every exception, so limit it to the transient ones; retrying a 400 "bad request" only wastes time.
Layer 3: shed load
chain.batch(inputs, config={"max_concurrency": 5})caps parallel calls.set_llm_cache(InMemoryCache())(fromlangchain_core.globalsandlangchain_core.caches) skips exact repeat calls.llm.with_fallbacks([backup_llm])switches to another model or provider when the main one keeps failing.
A real-life example
An e-commerce company in Bengaluru re-writes 60,000 product descriptions every night before a sale. Their account allows 500 requests per minute. The first run fired 50 parallel calls, hit 429s within 20 seconds, and the default retries piled on more requests, so the job took 5 hours and 7% of items failed. The fix: an InMemoryRateLimiter at 7 requests per second in the single worker process (about 420 per minute, under the limit), max_concurrency=20, with_retry only for RateLimitError and timeouts, and return_exceptions=True so failures were collected and re-queued. The job finished in about 2.5 hours with no failed items.
Follow-up questions to expect
- "How do you handle the tokens-per-minute limit?" — Estimate tokens per request, lower concurrency for long inputs, and use the provider's batch API for large offline jobs.
- "What is jitter and why add it?" — A small random delay added to each backoff, so many clients that failed together do not retry at the same moment.
- "What if one provider is down for an hour?" — Retries will not help;
with_fallbacksto a second provider or a smaller model keeps the service up.