LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug LangChain agent tool selection?


Two stock tools, before and after rewriting descriptionsBefore: 58% right first tool• get_stock_price: Gets stock data• get_company_info: Gets company data• Model guesses between near-twins• Price questions get a profileAfter: 95% right first tool• Price: current price for an NSE ticker• Info: static profile, never prices• Each says when not to use it• Same model, no prompt change
The tool description is the only thing the model reads when choosing, so that is where tool-selection bugs are fixed.

What you need to know

How an agent picks a tool

With tool calling, you send the model a list of tool schemas — name, description, JSON schema of the arguments. The model replies either with text or with one or more tool calls. The agent runs them and sends back ToolMessages with the results, and the loop repeats. So the only inputs to the choice are: the system prompt, the conversation, and the tool schemas.

Seeing the decisions

Python
from langchain.messages import AIMessage, ToolMessageresult = agent.invoke({"messages": [{"role": "user", "content": "Price of TCS?"}]})for m in result["messages"]:    if isinstance(m, AIMessage) and m.tool_calls:        for call in m.tool_calls:            print("CALL  ", call["name"], call["args"])    elif isinstance(m, ToolMessage):        print("RESULT", m.name, m.status, str(m.content)[:200])

To watch it live, stream updates — each chunk is keyed by the step that produced it ("model" or "tools"):

Python
for update in agent.stream({"messages": [...]}, stream_mode="updates"):    for step, data in update.items():        print(step, data["messages"][-1])

To see exactly what the model was shown, print tool.tool_call_schema.model_json_schema() for each tool, or open the model run in LangSmith, which displays the tool definitions sent with the request.

Legacy note: with the old AgentExecutor you set return_intermediate_steps=True and read result["intermediate_steps"] as (action, observation) pairs. That API now lives in langchain-classic; create_agent replaced it.

The usual causes and fixes

SymptomCauseFix
Picks the wrong one of two toolsOverlapping descriptionsMerge them, or state when to use each
Never calls a toolDescription too vague, or prompt allows answering from memorySay "always use X for current prices"
Wrong argumentsLoose types (str for a date)Typed schema: date, Literal[...], field descriptions
Calls tools in a loopTool errors hidden, or no stopping ruleReturn clear errors; ToolCallLimitMiddleware
Too many tools (30+)Choice gets harderGroup tools, use sub-agents, or LLMToolSelectorMiddleware to pre-select

Measure the fix

Build 30–100 queries with the expected tool. In LangSmith, an evaluator can check outputs["messages"] for the first tool call's name. Run it before and after each description change.

A real-life example

A finance assistant has two tools: get_stock_price(ticker) and get_company_info(name). Users asking "How is Infosys doing today?" got a company profile instead of the price 40% of the time.

The trace showed the model choosing get_company_info(name="Infosys"). The descriptions were "Gets stock data" and "Gets company data" — nearly the same. The engineer rewrote them: "Current share price and today's % change for an NSE ticker such as INFY. Use for any question about price, today, or performance," and "Static company profile: sector, founding year, CEO. Do not use for prices." They also added a Literal of allowed exchanges. On 60 labelled queries, correct first-tool choice went from 58% to 95% — with no model change.

Follow-up questions to expect

  • "Would a bigger model fix it?" — Sometimes, but it costs more on every call; fix descriptions and schemas first.
  • "How do you force a tool?" — Many providers support tool_choice (for example "any" or a named tool) when you bind tools; use it for steps that must call a tool.
  • "How do you stop tool loops?" — ToolCallLimitMiddleware or ModelCallLimitMiddleware, a recursion_limit you set yourself in the run config (LangGraph's default is very high — 10,007 in 1.2, it was 25 in early versions), and tool errors that tell the model what to do next.