Course Content
LangChain Mastery
7 sections · 109 lessons
How do you debug LangChain agent tool selection?
What you need to know
How an agent picks a tool
With tool calling, you send the model a list of tool schemas — name, description, JSON schema of the arguments. The model replies either with text or with one or more tool calls. The agent runs them and sends back ToolMessages with the results, and the loop repeats. So the only inputs to the choice are: the system prompt, the conversation, and the tool schemas.
Seeing the decisions
1from langchain.messages import AIMessage, ToolMessage23result = agent.invoke({"messages": [{"role": "user", "content": "Price of TCS?"}]})4for m in result["messages"]:5 if isinstance(m, AIMessage) and m.tool_calls:6 for call in m.tool_calls:7 print("CALL ", call["name"], call["args"])8 elif isinstance(m, ToolMessage):9 print("RESULT", m.name, m.status, str(m.content)[:200])To watch it live, stream updates — each chunk is keyed by the step that produced it ("model" or "tools"):
for update in agent.stream({"messages": [...]}, stream_mode="updates"): for step, data in update.items(): print(step, data["messages"][-1])To see exactly what the model was shown, print tool.tool_call_schema.model_json_schema() for each tool, or open the model run in LangSmith, which displays the tool definitions sent with the request.
Legacy note: with the old AgentExecutor you set return_intermediate_steps=True and read result["intermediate_steps"] as (action, observation) pairs. That API now lives in langchain-classic; create_agent replaced it.
The usual causes and fixes
| Symptom | Cause | Fix |
|---|---|---|
| Picks the wrong one of two tools | Overlapping descriptions | Merge them, or state when to use each |
| Never calls a tool | Description too vague, or prompt allows answering from memory | Say "always use X for current prices" |
| Wrong arguments | Loose types (str for a date) | Typed schema: date, Literal[...], field descriptions |
| Calls tools in a loop | Tool errors hidden, or no stopping rule | Return clear errors; ToolCallLimitMiddleware |
| Too many tools (30+) | Choice gets harder | Group tools, use sub-agents, or LLMToolSelectorMiddleware to pre-select |
Measure the fix
Build 30–100 queries with the expected tool. In LangSmith, an evaluator can check outputs["messages"] for the first tool call's name. Run it before and after each description change.
A real-life example
A finance assistant has two tools: get_stock_price(ticker) and get_company_info(name). Users asking "How is Infosys doing today?" got a company profile instead of the price 40% of the time.
The trace showed the model choosing get_company_info(name="Infosys"). The descriptions were "Gets stock data" and "Gets company data" — nearly the same. The engineer rewrote them: "Current share price and today's % change for an NSE ticker such as INFY. Use for any question about price, today, or performance," and "Static company profile: sector, founding year, CEO. Do not use for prices." They also added a Literal of allowed exchanges. On 60 labelled queries, correct first-tool choice went from 58% to 95% — with no model change.
Follow-up questions to expect
- "Would a bigger model fix it?" — Sometimes, but it costs more on every call; fix descriptions and schemas first.
- "How do you force a tool?" — Many providers support
tool_choice(for example "any" or a named tool) when you bind tools; use it for steps that must call a tool. - "How do you stop tool loops?" —
ToolCallLimitMiddlewareorModelCallLimitMiddleware, arecursion_limityou set yourself in the run config (LangGraph's default is very high — 10,007 in 1.2, it was 25 in early versions), and tool errors that tell the model what to do next.