Course Content
LangChain Mastery
7 sections · 109 lessons
How do you handle errors in LangChain chains?
What you need to know
First sort the error, because the right response depends on the kind.
| Error kind | Examples | Right response |
|---|---|---|
| Transient | Timeout, 429 rate limit, 500/503 | Retry with backoff |
| Provider down | Repeated 5xx, connection refused | Fall back to another model |
| Bad input | 400, context too long, content filtered | Fix or shorten input; do not retry as-is |
| Bad output | Parse or validation failure | One retry with the error, then review |
| Your code | KeyError in a function step | Fix the bug |
Retry
1import openai23safe = chain.with_retry(4 retry_if_exception_type=(openai.RateLimitError, openai.APITimeoutError,5 openai.InternalServerError),6 wait_exponential_jitter=True,7 stop_after_attempt=3,8)By default with_retry retries on any exception, which is why the types are listed. Also set timeout on the model, or a hung call never becomes an error to retry.
Fall back
1robust = primary_chain.with_fallbacks(2 [backup_chain, RunnableLambda(lambda x: "Sorry, please try again in a minute.")],3 exception_key="error",4)Fallbacks run in order, each only if the one before raised. With exception_key, the fallback receives the error in its input dict so it can log it or change its behaviour. A fallback can be a whole chain with its own prompt, not only another model.
Catch what's yours
1from langchain_core.exceptions import OutputParserException2from pydantic import ValidationError34try:5 result = robust.invoke(payload)6except (OutputParserException, ValidationError) as e:7 log.warning("bad output: %s", e)8 result = send_to_review(payload)In batch jobs, return_exceptions=True keeps one failure from stopping the rest.
A real-life example
A support bot over a telecom company's help-centre docs had one provider outage lasting 40 minutes during an IPL final, the busiest evening of the month. With only retries, every request waited through three attempts, about 30 seconds, then failed. After that night the team added with_fallbacks to a second provider with a prompt tuned for it, and a final fallback that returns "We're having trouble; here are the top help articles" with links from retrieval alone. In the next outage, users saw answers from the backup model within 3 seconds, and the dashboard showed the fallback rate rising, which alerted the on-call engineer.
Follow-up questions to expect
- "Retry at the model level or the chain level?" — Model-level
max_retriescovers provider errors cheaply; chain-levelwith_retryalso covers retrieval and your own steps. Don't stack both with high counts, or attempts multiply. - "How do you handle a context-length error?" — Catch it and shorten the input (fewer documents, trimmed history), or fall back to a model with a larger context window.
- "How do you know fallbacks are firing?" — Log or tag fallback runs and alert on the rate.