Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to retry LangChain chain execution on failure.
What you need to know
The function
1import openai2from langchain_core.runnables import Runnable34TRANSIENT = (openai.RateLimitError, openai.APITimeoutError,5 openai.APIConnectionError, openai.InternalServerError)67def make_resilient(chain: Runnable, backup: Runnable | None = None,8 attempts: int = 3) -> Runnable:9 robust = chain.with_retry(10 retry_if_exception_type=TRANSIENT,11 stop_after_attempt=attempts,12 wait_exponential_jitter=True, # waits roughly 1s, 2s, 4s... plus random jitter13 )14 return robust.with_fallbacks([backup]) if backup else robust1516answer_chain = make_resilient(prompt | primary_llm | parser,17 backup=prompt | backup_llm | parser)retry_if_exception_type— only these exceptions trigger a retry. Anything else fails at once.stop_after_attempt— total attempts, including the first.wait_exponential_jitter— waits grow exponentially with random jitter, so thousands of clients don't retry at the same instant.with_fallbacks— tries the next runnable if the first still fails. The fallback should include the prompt and parser, since prompts sometimes need to differ per model.
Two layers of retry
Provider SDKs already retry a little (the OpenAI client retries twice by default; set max_retries on ChatOpenAI). If you also add with_retry(stop_after_attempt=3), one call can become 3 × 3 = 9 requests. Decide which layer owns retries.
What to retry and what not
| Retry | Do not retry |
|---|---|
| 429 rate limit | 400 bad request, context too long |
| Timeouts, connection errors | 401/403 auth errors |
| 500/502/503 | Content-policy refusals |
| An occasional parse failure (once, maybe with a repair prompt) | Your own KeyError or validation errors |
Idempotency
Retrying a model call is safe. Retrying a tool that has side effects — booking a ticket, filing leave, sending money — can do it twice. Such tools need an idempotency key (a unique id per request that the backend uses to ignore duplicates) before you add retries.
In agents
1from langchain.agents import create_agent2from langchain.agents.middleware import ModelRetryMiddleware, ModelFallbackMiddleware34agent = create_agent("openai:gpt-5.4-mini", tools=[...], middleware=[5 ModelRetryMiddleware(max_retries=2),6 ModelFallbackMiddleware("anthropic:claude-sonnet-4-6"),7])A real-life example
A travel-booking agent calls a model and a flight-search API. During a sale, the model provider returns 429 for about 1% of calls. Without retries, those users got errors. With a plain for loop retrying immediately, 2,000 workers retried at the same moment and made the rate limiting worse.
With with_retry (3 attempts, exponential jitter) plus a fallback to a second provider, failed requests dropped to near zero; p99 latency rose by about 2 seconds during the spike. Separately, the book_flight tool was not retried automatically — a retry after a timeout had once created two bookings. The team added an idempotency key and then enabled a single retry.
Follow-up questions to expect
- "Why jitter?" — Without it, clients that failed together retry together, causing a "thundering herd" that repeats the overload.
- "Where do you put the retry in a chain?" — Around the smallest step that can fail (usually the model call), so a retry does not repeat retrieval or other work.
- "How do you know retries are happening?" — Traces show each attempt; also count retries in metrics, since a rising rate is an early warning.