LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you handle errors in LangChain chains?


What a failing request falls throughRetry — timeouts, 429, 5xx only, with backoffFallback chain — second provider, its own promptCanned fallback — top help articles, no modelCatch parse errors — retry once, then review
Sorting the error first matters: a too-long prompt fails the same way on every retry, so it needs a shorter input, not another attempt.

What you need to know

First sort the error, because the right response depends on the kind.

Error kindExamplesRight response
TransientTimeout, 429 rate limit, 500/503Retry with backoff
Provider downRepeated 5xx, connection refusedFall back to another model
Bad input400, context too long, content filteredFix or shorten input; do not retry as-is
Bad outputParse or validation failureOne retry with the error, then review
Your codeKeyError in a function stepFix the bug

Retry

Python
import openaisafe = chain.with_retry(    retry_if_exception_type=(openai.RateLimitError, openai.APITimeoutError,                             openai.InternalServerError),    wait_exponential_jitter=True,    stop_after_attempt=3,)

By default with_retry retries on any exception, which is why the types are listed. Also set timeout on the model, or a hung call never becomes an error to retry.

Fall back

Python
robust = primary_chain.with_fallbacks(    [backup_chain, RunnableLambda(lambda x: "Sorry, please try again in a minute.")],    exception_key="error",)

Fallbacks run in order, each only if the one before raised. With exception_key, the fallback receives the error in its input dict so it can log it or change its behaviour. A fallback can be a whole chain with its own prompt, not only another model.

Catch what's yours

Python
from langchain_core.exceptions import OutputParserExceptionfrom pydantic import ValidationErrortry:    result = robust.invoke(payload)except (OutputParserException, ValidationError) as e:    log.warning("bad output: %s", e)    result = send_to_review(payload)

In batch jobs, return_exceptions=True keeps one failure from stopping the rest.

A real-life example

A support bot over a telecom company's help-centre docs had one provider outage lasting 40 minutes during an IPL final, the busiest evening of the month. With only retries, every request waited through three attempts, about 30 seconds, then failed. After that night the team added with_fallbacks to a second provider with a prompt tuned for it, and a final fallback that returns "We're having trouble; here are the top help articles" with links from retrieval alone. In the next outage, users saw answers from the backup model within 3 seconds, and the dashboard showed the fallback rate rising, which alerted the on-call engineer.

Follow-up questions to expect

  • "Retry at the model level or the chain level?" — Model-level max_retries covers provider errors cheaply; chain-level with_retry also covers retrieval and your own steps. Don't stack both with high counts, or attempts multiply.
  • "How do you handle a context-length error?" — Catch it and shorten the input (fewer documents, trimmed history), or fall back to a model with a larger context window.
  • "How do you know fallbacks are firing?" — Log or tag fallback runs and alert on the rate.