LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to retry LangChain chain execution on failure.


One call under a rate limit, with with_retryattempt 1: 429wait about 1splus jitterattempt 2: 429wait about 2splus jitterattempt 3: 200 OKnulljitterspreads clientsstop_after_attempt=3A 400 or a validation error is never retried — it would fail the same way.
Exponential waits with jitter stop thousands of workers from retrying in the same instant and recreating the overload.

What you need to know

The function

Python
import openaifrom langchain_core.runnables import RunnableTRANSIENT = (openai.RateLimitError, openai.APITimeoutError,             openai.APIConnectionError, openai.InternalServerError)def make_resilient(chain: Runnable, backup: Runnable | None = None,                   attempts: int = 3) -> Runnable:    robust = chain.with_retry(        retry_if_exception_type=TRANSIENT,        stop_after_attempt=attempts,        wait_exponential_jitter=True,      # waits roughly 1s, 2s, 4s... plus random jitter    )    return robust.with_fallbacks([backup]) if backup else robustanswer_chain = make_resilient(prompt | primary_llm | parser,                              backup=prompt | backup_llm | parser)
  • retry_if_exception_type — only these exceptions trigger a retry. Anything else fails at once.
  • stop_after_attempt — total attempts, including the first.
  • wait_exponential_jitter — waits grow exponentially with random jitter, so thousands of clients don't retry at the same instant.
  • with_fallbacks — tries the next runnable if the first still fails. The fallback should include the prompt and parser, since prompts sometimes need to differ per model.

Two layers of retry

Provider SDKs already retry a little (the OpenAI client retries twice by default; set max_retries on ChatOpenAI). If you also add with_retry(stop_after_attempt=3), one call can become 3 × 3 = 9 requests. Decide which layer owns retries.

What to retry and what not

RetryDo not retry
429 rate limit400 bad request, context too long
Timeouts, connection errors401/403 auth errors
500/502/503Content-policy refusals
An occasional parse failure (once, maybe with a repair prompt)Your own KeyError or validation errors

Idempotency

Retrying a model call is safe. Retrying a tool that has side effects — booking a ticket, filing leave, sending money — can do it twice. Such tools need an idempotency key (a unique id per request that the backend uses to ignore duplicates) before you add retries.

In agents

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import ModelRetryMiddleware, ModelFallbackMiddlewareagent = create_agent("openai:gpt-5.4-mini", tools=[...], middleware=[    ModelRetryMiddleware(max_retries=2),    ModelFallbackMiddleware("anthropic:claude-sonnet-4-6"),])

A real-life example

A travel-booking agent calls a model and a flight-search API. During a sale, the model provider returns 429 for about 1% of calls. Without retries, those users got errors. With a plain for loop retrying immediately, 2,000 workers retried at the same moment and made the rate limiting worse.

With with_retry (3 attempts, exponential jitter) plus a fallback to a second provider, failed requests dropped to near zero; p99 latency rose by about 2 seconds during the spike. Separately, the book_flight tool was not retried automatically — a retry after a timeout had once created two bookings. The team added an idempotency key and then enabled a single retry.

Follow-up questions to expect

  • "Why jitter?" — Without it, clients that failed together retry together, causing a "thundering herd" that repeats the overload.
  • "Where do you put the retry in a chain?" — Around the smallest step that can fail (usually the model call), so a retry does not repeat retrieval or other work.
  • "How do you know retries are happening?" — Traces show each attempt; also count retries in metrics, since a rising rate is an early warning.