Course Content
LangChain Mastery
7 sections · 109 lessons
How do you handle tool failures in LangChain agents?
What you need to know
Three kinds of failure
| Kind | Example | Best handler |
|---|---|---|
| Bad arguments | pincode="56" | Schema validation; the error goes back to the model to fix |
| Transient | timeout, HTTP 503 | Retry with backoff inside the runtime, not via the model |
| Permanent | 404 order not found, service off | Tell the model clearly so it can explain or try another tool |
The main idea: a failed tool should usually produce an observation, not a crash. The model can then ask the user for a correct order ID, pick another tool, or apologise honestly.
Default behaviour in create_agent
In LangChain 1.x, argument validation errors are caught by the tool node and returned to the model automatically. Exceptions raised while the tool runs are not caught by default — they propagate and end the run. You opt in to catching them.
Option 1: ToolException in the tool
1import httpx2from langchain_core.tools import tool, ToolException34@tool5def get_order(order_id: str) -> str:6 """Look up an order's status by ID, e.g. 'ORD-1042'."""7 r = httpx.get(f"{ORDERS_API}/{order_id}", timeout=5) # network errors propagate8 if r.status_code == 404:9 raise ToolException(f"No order {order_id}. Ask the user to re-check the ID.")10 r.raise_for_status()11 return r.json()["status"]1213get_order.handle_tool_error = True # ToolException text becomes the observationhandle_tool_error accepts True, a fixed string, or a function that builds the message from the exception. It only catches ToolException, so real bugs (a KeyError in your code) still surface. Timeouts and 5xx errors are deliberately left to raise, so the retry middleware below can retry them.
Option 2: middleware for all tools
1from langchain.agents import create_agent2from langchain.agents.middleware import ToolRetryMiddleware, wrap_tool_call3from langchain.messages import ToolMessage45@wrap_tool_call6def errors_to_messages(request, handler):7 try:8 return handler(request)9 except Exception as e:10 name = request.tool_call["name"]11 log.warning("tool %s failed: %r", name, e) # full detail in logs12 return ToolMessage(content=f"{name} is unavailable right now. "13 "Tell the user and do not call it again.",14 tool_call_id=request.tool_call["id"], status="error")1516def transient(e: Exception) -> bool:17 return isinstance(e, httpx.TransportError) or (18 isinstance(e, httpx.HTTPStatusError) and e.response.status_code >= 500)1920agent = create_agent(model, tools=[get_order], middleware=[21 errors_to_messages, # outer: catches the rest22 ToolRetryMiddleware(max_retries=2, retry_on=transient), # inner: retries blips23])Middleware listed first wraps the ones after it. ToolRetryMiddleware retries transient errors with exponential backoff and jitter before the model ever sees the failure — cheaper than spending another model call. The outer wrap_tool_call function catches whatever is left and turns it into a ToolMessage. Newer 1.x releases also ship a built-in ToolErrorMiddleware for the same job.
Guard rails around it
- A
timeouton every network call; a hanging tool blocks the whole run. ToolCallLimitMiddleware(tool_name="get_order", run_limit=3)so a broken tool is not called 20 times.- Log every failure with tool name and arguments. Repeated failures with odd arguments usually mean a vague tool description, not a flaky service.
A real-life example
A help-centre support bot calls a billing API that returns HTTP 503 about 2% of the time during month-end load. Before the fix, each 503 raised an exception, the user saw "Something went wrong", and around 400 conversations a day ended there.
The team adds ToolRetryMiddleware(max_retries=2) for transport errors and 5xx responses, which absorbs most blips in under 3 seconds. When retries run out, the outer middleware returns "Billing is temporarily unavailable" as the tool result, and the model tells the customer it cannot see the bill right now and offers a callback. Hard failures drop to about 15 a day, and none show a raw error.
Follow-up questions to expect
- "Should retries happen in the tool or through the model?" — Transient retries belong in code (cheap, fast); only let the model retry when it needs to change the arguments.
- "What does the model see when a tool fails?" — A
ToolMessagewith the error text, often withstatus="error", which it reads like any other result. - "How do you avoid leaking internals in the error?" — Write the message for the model: what failed and what to do next, without stack traces, hostnames or keys.