Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 2: Incorrect Tool Selection
Scenario: a LangChain agent has a dozen or more tools and keeps invoking the wrong one. How do you fix tool selection?
What you need to know
When LangChain binds tools to a chat model, it converts each tool into a name, a description (from the docstring) and a JSON Schema (from the type hints). That is all the model knows. If two descriptions sound alike, the model has no basis to choose correctly.
Rewrite the contract
1from typing import Literal2from langchain_core.tools import tool34@tool5def search_orders(customer_id: str,6 status: Literal["pending", "shipped", "delivered"] | None = None) -> list[dict]:7 """Look up a customer's ORDERS by customer_id: order status, delivery dates, order history.8 Do NOT use for refunds or payments; use `search_payments` for those."""9 return orders_api.search(customer_id=customer_id, status=status)Three things changed from a typical first draft:
- What it is for, in the words users use ("delivery dates", "order history").
- What it is not for, naming the tool it gets confused with.
Literalinstead ofstrforstatus, so the model cannot invent"in transit".
Structural fixes, in order
- Fewer tools per call — past roughly a dozen or two overlapping tools, selection suffers. Retrieve the relevant subset by embedding tool descriptions, or split into sub-agents by domain under a supervisor.
- Sharper schemas — enums and required fields prevent both wrong calls and invalid arguments.
- Routing examples — three to five real "query goes to tool" pairs from the misfires, in the system prompt.
- Forced choice where the step is known —
bind_tools(tools, tool_choice="search_orders")when the flow already knows which tool must run.
Find the confusable pairs
From the traces, build a confusion table: for each wrong call, which tool was right? Misfires cluster in a few pairs.
| Called | Should have called | Count |
|---|---|---|
search_orders | search_payments | 41 |
get_customer | get_account_settings | 18 |
create_ticket | update_ticket | 11 |
Fix the top pairs first; they usually account for most of the errors.
Make it measurable
Create a LangSmith dataset of real queries labelled with the correct tool, and an evaluator that checks the first tool call. Run it on every change to a tool, prompt or model. Without the number, "better descriptions" is guesswork; with it, each rewrite is verified in minutes.
A real-life example
Scenario (illustrative numbers). A fintech's support agent has 16 tools. On a 300-query dataset built from traces, it picks the right tool 74% of the time. The confusion table shows search_orders versus search_payments and get_customer versus get_account_settings account for 60% of errors.
Rewriting those four descriptions with "do not use for" lines and switching two str arguments to Literal lifts accuracy to 87%. Splitting the tools into two sub-agents, "orders and delivery" and "payments and account", under a supervisor lifts it to 94%. The dataset now runs in CI, and a later docstring edit that dropped a "do not use" line is caught before release.
Follow-up questions to expect
- "Would a bigger model fix it?" — Sometimes partly, at higher cost per call. Clear descriptions help every model and cost nothing to run.
- "How do you evaluate multi-step tool use?" — Score the whole trajectory against an expected sequence, or at least the first call plus the final answer.
- "Where do the descriptions come from?" — The docstring and type hints by default; you can also pass
description=explicitly to@tool.