Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 2: Incorrect Tool Selection


Scenario: a LangChain agent has a dozen or more tools and keeps invoking the wrong one. How do you fix tool selection?

What you need to know

When LangChain binds tools to a chat model, it converts each tool into a name, a description (from the docstring) and a JSON Schema (from the type hints). That is all the model knows. If two descriptions sound alike, the model has no basis to choose correctly.

Rewrite the contract

Python
from typing import Literalfrom langchain_core.tools import tool@tooldef search_orders(customer_id: str,                  status: Literal["pending", "shipped", "delivered"] | None = None) -> list[dict]:    """Look up a customer's ORDERS by customer_id: order status, delivery dates, order history.    Do NOT use for refunds or payments; use `search_payments` for those."""    return orders_api.search(customer_id=customer_id, status=status)

Three things changed from a typical first draft:

  • What it is for, in the words users use ("delivery dates", "order history").
  • What it is not for, naming the tool it gets confused with.
  • Literal instead of str for status, so the model cannot invent "in transit".

Structural fixes, in order

  1. Fewer tools per call — past roughly a dozen or two overlapping tools, selection suffers. Retrieve the relevant subset by embedding tool descriptions, or split into sub-agents by domain under a supervisor.
  2. Sharper schemas — enums and required fields prevent both wrong calls and invalid arguments.
  3. Routing examples — three to five real "query goes to tool" pairs from the misfires, in the system prompt.
  4. Forced choice where the step is known — bind_tools(tools, tool_choice="search_orders") when the flow already knows which tool must run.

Find the confusable pairs

From the traces, build a confusion table: for each wrong call, which tool was right? Misfires cluster in a few pairs.

CalledShould have calledCount
search_orderssearch_payments41
get_customerget_account_settings18
create_ticketupdate_ticket11

Fix the top pairs first; they usually account for most of the errors.

Make it measurable

Create a LangSmith dataset of real queries labelled with the correct tool, and an evaluator that checks the first tool call. Run it on every change to a tool, prompt or model. Without the number, "better descriptions" is guesswork; with it, each rewrite is verified in minutes.

A real-life example

Scenario (illustrative numbers). A fintech's support agent has 16 tools. On a 300-query dataset built from traces, it picks the right tool 74% of the time. The confusion table shows search_orders versus search_payments and get_customer versus get_account_settings account for 60% of errors.

Rewriting those four descriptions with "do not use for" lines and switching two str arguments to Literal lifts accuracy to 87%. Splitting the tools into two sub-agents, "orders and delivery" and "payments and account", under a supervisor lifts it to 94%. The dataset now runs in CI, and a later docstring edit that dropped a "do not use" line is caught before release.

Follow-up questions to expect

  • "Would a bigger model fix it?" — Sometimes partly, at higher cost per call. Clear descriptions help every model and cost nothing to run.
  • "How do you evaluate multi-step tool use?" — Score the whole trajectory against an expected sequence, or at least the first call plus the final answer.
  • "Where do the descriptions come from?" — The docstring and type hints by default; you can also pass description= explicitly to @tool.