Course Content
LangGraph Agents
7 sections · 49 lessons
How do you implement fallback paths (tool failure → alternate tool → safe response)?
What you need to know
| Failure class | Example | Response |
|---|---|---|
| Transient | 503, 429, timeout | Retry with backoff at the node |
| Bad input | Invalid order id, schema error | Repair: send the error back to the model |
| Unavailable | Provider down, quota gone | Alternate tool or cached data |
| Exhausted | All options failed | Safe response, maybe a ticket |
Retry policy — know the default
1import httpx2from langgraph.types import RetryPolicy34def retryable(e: Exception) -> bool:5 if isinstance(e, httpx.HTTPStatusError):6 return e.response.status_code in (429, 500, 502, 503, 504)7 return isinstance(e, (httpx.TimeoutException, ConnectionError, TimeoutError))89builder.add_node("search", search,10 retry_policy=RetryPolicy(max_attempts=3, initial_interval=1.0,11 retry_on=retryable))The default retry_on retries ConnectionError and HTTP 5xx, but not 4xx (so not 429) and not OSError subclasses such as Python's TimeoutError. Write your own predicate.
Routing after the retries
1def search(state):2 try:3 return {"results": primary_search(state["query"]), "error": None}4 except ProviderDown as e:5 return {"error": f"primary: {e}", "tried": state["tried"] + ["primary"]}67def after_search(state) -> str:8 if state["error"] is None:9 return "answer"10 if "backup" not in state["tried"]:11 return "backup_search"12 return "safe_response"LangGraph 1.2 also accepts error_handler= on add_node: a function (state, error: NodeError) that runs after retries are exhausted and can return Command(goto="backup_search"). It keeps the try/except out of the node.
A real-life example
A research agent uses a paid web-search API, with a free news index as backup. In one week the paid API returned 503 for 40 minutes and 429 during a traffic spike. With the right policy: the 503s were retried (three attempts, 1 s, 2 s, 4 s) and most recovered; the 429s were retried with the custom predicate; during the 40-minute outage, 312 runs moved to backup_search and answered with a note "results from news sources only"; 9 runs where both failed reached safe_response: "I couldn't search right now; I've saved your question and will email you the report." No run crashed and no answer was made up.
Follow-up questions to expect
- "Where do retries live — the HTTP client or the graph?" — One place only. Retries at both layers multiply: 3 times 3 is 9 calls to a struggling service.
- "How do you stop fallback loops?" — Track which tools were tried in state and never route back to one already tried.
- "What does
ToolNodedo on errors?" — By default it turns argument-validation errors into aToolMessagethe model can read, and re-raises other exceptions. Sethandle_tool_errorsto change that.