LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

How do you implement fallback paths (tool failure → alternate tool → safe response)?


Match the recovery to the failureTransient 503 or 429 — RetryPolicy with retry_onBad arguments — repair with the errorProvider down — backup_search, cached dataAll failed — safe_response and a ticket
The default retry predicate skips 429s and timeouts, so a policy that looks complete can fail on the most common error.

What you need to know

Failure classExampleResponse
Transient503, 429, timeoutRetry with backoff at the node
Bad inputInvalid order id, schema errorRepair: send the error back to the model
UnavailableProvider down, quota goneAlternate tool or cached data
ExhaustedAll options failedSafe response, maybe a ticket

Retry policy — know the default

Python
import httpxfrom langgraph.types import RetryPolicydef retryable(e: Exception) -> bool:    if isinstance(e, httpx.HTTPStatusError):        return e.response.status_code in (429, 500, 502, 503, 504)    return isinstance(e, (httpx.TimeoutException, ConnectionError, TimeoutError))builder.add_node("search", search,                 retry_policy=RetryPolicy(max_attempts=3, initial_interval=1.0,                                          retry_on=retryable))

The default retry_on retries ConnectionError and HTTP 5xx, but not 4xx (so not 429) and not OSError subclasses such as Python's TimeoutError. Write your own predicate.

Routing after the retries

Python
def search(state):    try:        return {"results": primary_search(state["query"]), "error": None}    except ProviderDown as e:        return {"error": f"primary: {e}", "tried": state["tried"] + ["primary"]}def after_search(state) -> str:    if state["error"] is None:        return "answer"    if "backup" not in state["tried"]:        return "backup_search"    return "safe_response"

LangGraph 1.2 also accepts error_handler= on add_node: a function (state, error: NodeError) that runs after retries are exhausted and can return Command(goto="backup_search"). It keeps the try/except out of the node.

A real-life example

A research agent uses a paid web-search API, with a free news index as backup. In one week the paid API returned 503 for 40 minutes and 429 during a traffic spike. With the right policy: the 503s were retried (three attempts, 1 s, 2 s, 4 s) and most recovered; the 429s were retried with the custom predicate; during the 40-minute outage, 312 runs moved to backup_search and answered with a note "results from news sources only"; 9 runs where both failed reached safe_response: "I couldn't search right now; I've saved your question and will email you the report." No run crashed and no answer was made up.

Follow-up questions to expect

  • "Where do retries live — the HTTP client or the graph?" — One place only. Retries at both layers multiply: 3 times 3 is 9 calls to a struggling service.
  • "How do you stop fallback loops?" — Track which tools were tried in state and never route back to one already tried.
  • "What does ToolNode do on errors?" — By default it turns argument-validation errors into a ToolMessage the model can read, and re-raises other exceptions. Set handle_tool_errors to change that.