LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you handle tool failures in LangChain agents?


Where each kind of tool failure is caughtBad arguments:schema errorback to modelTransient 503or timeout:retried in codeRetries usedup: errorToolMessageModel explains orpicks another toolCall limitstops adead toolMonth-end 503s dropped from 400 failed chats a day to about 15.
A failure should become an observation the model can act on, and only the retry loop in code should spend time on blips.

What you need to know

Three kinds of failure

KindExampleBest handler
Bad argumentspincode="56"Schema validation; the error goes back to the model to fix
Transienttimeout, HTTP 503Retry with backoff inside the runtime, not via the model
Permanent404 order not found, service offTell the model clearly so it can explain or try another tool

The main idea: a failed tool should usually produce an observation, not a crash. The model can then ask the user for a correct order ID, pick another tool, or apologise honestly.

Default behaviour in create_agent

In LangChain 1.x, argument validation errors are caught by the tool node and returned to the model automatically. Exceptions raised while the tool runs are not caught by default — they propagate and end the run. You opt in to catching them.

Option 1: ToolException in the tool

Python
import httpxfrom langchain_core.tools import tool, ToolException@tooldef get_order(order_id: str) -> str:    """Look up an order's status by ID, e.g. 'ORD-1042'."""    r = httpx.get(f"{ORDERS_API}/{order_id}", timeout=5)  # network errors propagate    if r.status_code == 404:        raise ToolException(f"No order {order_id}. Ask the user to re-check the ID.")    r.raise_for_status()    return r.json()["status"]get_order.handle_tool_error = True   # ToolException text becomes the observation

handle_tool_error accepts True, a fixed string, or a function that builds the message from the exception. It only catches ToolException, so real bugs (a KeyError in your code) still surface. Timeouts and 5xx errors are deliberately left to raise, so the retry middleware below can retry them.

Option 2: middleware for all tools

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import ToolRetryMiddleware, wrap_tool_callfrom langchain.messages import ToolMessage@wrap_tool_calldef errors_to_messages(request, handler):    try:        return handler(request)    except Exception as e:        name = request.tool_call["name"]        log.warning("tool %s failed: %r", name, e)          # full detail in logs        return ToolMessage(content=f"{name} is unavailable right now. "                                   "Tell the user and do not call it again.",                           tool_call_id=request.tool_call["id"], status="error")def transient(e: Exception) -> bool:    return isinstance(e, httpx.TransportError) or (        isinstance(e, httpx.HTTPStatusError) and e.response.status_code >= 500)agent = create_agent(model, tools=[get_order], middleware=[    errors_to_messages,                                   # outer: catches the rest    ToolRetryMiddleware(max_retries=2, retry_on=transient),  # inner: retries blips])

Middleware listed first wraps the ones after it. ToolRetryMiddleware retries transient errors with exponential backoff and jitter before the model ever sees the failure — cheaper than spending another model call. The outer wrap_tool_call function catches whatever is left and turns it into a ToolMessage. Newer 1.x releases also ship a built-in ToolErrorMiddleware for the same job.

Guard rails around it

  • A timeout on every network call; a hanging tool blocks the whole run.
  • ToolCallLimitMiddleware(tool_name="get_order", run_limit=3) so a broken tool is not called 20 times.
  • Log every failure with tool name and arguments. Repeated failures with odd arguments usually mean a vague tool description, not a flaky service.

A real-life example

A help-centre support bot calls a billing API that returns HTTP 503 about 2% of the time during month-end load. Before the fix, each 503 raised an exception, the user saw "Something went wrong", and around 400 conversations a day ended there.

The team adds ToolRetryMiddleware(max_retries=2) for transport errors and 5xx responses, which absorbs most blips in under 3 seconds. When retries run out, the outer middleware returns "Billing is temporarily unavailable" as the tool result, and the model tells the customer it cannot see the bill right now and offers a callback. Hard failures drop to about 15 a day, and none show a raw error.

Follow-up questions to expect

  • "Should retries happen in the tool or through the model?" — Transient retries belong in code (cheap, fast); only let the model retry when it needs to change the arguments.
  • "What does the model see when a tool fails?" — A ToolMessage with the error text, often with status="error", which it reads like any other result.
  • "How do you avoid leaking internals in the error?" — Write the message for the model: what failed and what to do next, without stack traces, hostnames or keys.