LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

How do you handle tool errors and recover (retry, alternate tool, partial results)?


What you need to know

What ToolNode does by default (1.x)

  • Invalid arguments (the model sent "many" for an int) — returned to the model as a ToolMessage describing the error.
  • Unknown tool name — returned as an error message listing valid tools.
  • Any other exception (a timeout, a database error) — re-raised, so the node fails. This changed in 1.0; older versions caught everything.

To change it, set handle_tool_errors:

Python
from langgraph.prebuilt import ToolNodedef explain(e: Exception) -> str:    return f"Tool failed: {e}. Check the order id, or use search_orders instead."tools_node = ToolNode(tools, handle_tool_errors=explain)          # callable# ToolNode(tools, handle_tool_errors=True)                        # catch all, generic text# ToolNode(tools, handle_tool_errors=(ConnectionError,))           # catch these types only

A plain string is returned as-is; it does not fill in {error}, so use a callable when you want the error text.

The layers

  1. Self-correction — the model sees the error and retries with better arguments. Cap it: two failures on the same tool usually means it is stuck.
  2. Retry transient errors — RetryPolicy(retry_on=...) on the tool node, or ToolRetryMiddleware(max_retries=2) in create_agent. No tokens spent.
  3. Alternate path — an edge or error_handler routes to a backup tool or cache.
  4. Partial results — in parallel work, return {"failed": [doc_id]} and continue.
  5. Honest answer — if nothing works, say what could not be done.

ModelFallbackMiddleware does the same for the model itself: if the primary model errors, it tries the next one.

A real-life example

A research agent that writes company profiles calls get_financials, web_search and get_news. Over a month: 4% of calls had bad arguments (a ticker instead of a company id) — fixed by self-correction on the next turn; 1.5% hit 503 or timeouts — 90% recovered on retry; one 2-hour outage of the financials provider — runs switched to cached quarterly data and the report said "figures as of last quarter". Before this layering, the agent turned the empty error string into "the company reported no revenue", which a client noticed.

Follow-up questions to expect

  • "Should the model see stack traces?" — No. Give a short, actionable message; long traces waste tokens and may leak internals.
  • "How do you know recovery is working?" — Track error rate per tool, retries per run, and the share of runs that end in the honest-failure path.
  • "Retry in the tool or in the graph?" — One layer only. The HTTP client and the graph both retrying three times means nine calls.