LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to handle tool timeouts in LangChain agents.


Three timeouts, each catching a different hangWhole run: 30 s, plus a model-call limitTool wrapper: asyncio.wait_for, 8 sHTTP client: 5 s connect and read
The client timeout catches a slow server, the wrapper catches what ignores it, and the run limit protects the user from many slow steps.

What you need to know

Why one timeout is not enough

LevelExampleCatches
Clienthttpx.AsyncClient(timeout=5), SQL statement_timeoutA slow server or query
Tool wrapperasyncio.wait_for(..., timeout=8)Retries that add up, DNS hangs, libraries without timeouts
Whole runasyncio.wait_for(agent.ainvoke(...), 30) + call limitMany slow steps, loops

Also set a timeout on the chat model itself (timeout= and max_retries= on the model class or init_chat_model), since a slow model call is just as bad as a slow tool.

The code

Python
import asyncio, httpxfrom langchain_core.tools import StructuredTool, ToolExceptionclient = httpx.AsyncClient(timeout=5.0)                      # level 1async def _lookup(query: str) -> str:    r = await client.get(CATALOGUE_URL, params={"q": query})    r.raise_for_status()    return r.text[:2000]async def lookup(query: str) -> str:    try:        return await asyncio.wait_for(_lookup(query), timeout=8)   # level 2    except asyncio.TimeoutError:        raise ToolException("Catalogue search timed out after 8s. Answer without it "                            "and tell the user live stock could not be checked.")catalogue_search = StructuredTool.from_function(    coroutine=lookup, name="catalogue_search",    description="Search the product catalogue by name or feature.",    handle_tool_error=True)async def answer(agent, question: str) -> str:    try:        out = await asyncio.wait_for(                                # level 3            agent.ainvoke({"messages": [{"role": "user", "content": question}]}),            timeout=30)        return out["messages"][-1].content    except asyncio.TimeoutError:        return "Sorry, this is taking too long. Please try again in a minute."

Walking through it

  • httpx.AsyncClient(timeout=5.0) applies to connect, read and write. Without it, httpx defaults to 5 seconds too, but many other libraries default to no timeout.
  • asyncio.wait_for cancels the coroutine after 8 seconds. It raises asyncio.TimeoutError (an alias of the built-in TimeoutError since Python 3.11).
  • ToolException + handle_tool_error=True turn the timeout into an observation. The model then answers from what it has, instead of the run crashing.
  • The run-level limit protects your API's own response time. Pair it with ModelCallLimitMiddleware so a loop of fast steps is also stopped.

Sync tools are different

asyncio.wait_for cannot interrupt blocking code such as requests.get or a sync database driver running in the event loop. Options: switch to an async client; run the sync call with await asyncio.to_thread(fn, ...) inside wait_for (the caller stops waiting, though the thread still finishes in the background); or use the library's own timeout. For sync tools in a sync agent, rely on the client timeout.

Choosing the numbers

Work backwards from the user. If a chat reply must arrive within 30 seconds and a typical run has 3 model calls of about 3 seconds each, tools can use about 20 seconds in total, so 5 to 8 seconds per tool call is reasonable. Log timeouts per tool; a tool that times out on more than 1% of calls needs fixing at the source.

A real-life example

An electronics store's product Q&A assistant calls a supplier stock API for items not held in the warehouse. The supplier API usually answers in 400 ms, but during sale events it sometimes hangs for over 60 seconds. The old tool had no timeout, so chats froze and the load balancer killed requests at 60 seconds with a blank error.

After the change, the client times out at 5 seconds, the wrapper at 8, and the run at 30. During the next sale, 4% of supplier calls timed out; in each case the assistant replied "I can't check the supplier's live stock right now; it usually ships in 5 to 7 days" within about 12 seconds. Blank errors dropped to zero.

Follow-up questions to expect

  • "What happens to the tool's work after a timeout?" — The coroutine is cancelled; a thread started with to_thread keeps running, so make tools idempotent and avoid timing out write operations halfway.
  • "Should you retry after a timeout?" — At most once, with backoff, and only for read-only calls; retrying a slow service under load makes it slower.
  • "How do you time out a write, like placing an order?" — Use an idempotency key so a retry cannot create a second order, and confirm the outcome before telling the user.