Advanced Prompting and Reasoning

Tool-Augmented Reasoning


A support agent has five tools. A customer types: "Where is my order #4471?"

The agent calls search_knowledge_base("where is my order") and returns a help article titled "Understanding Shipping Times". The order status was one lookup_order call away and the agent never made it.

Nobody wrote a bug. The tool descriptions looked like this:

Text
lookup_order          "Look up an order."search_orders         "Search orders."search_customers      "Search customers."search_tickets        "Search support tickets."search_knowledge_base "Search the knowledge base for answers to                       customer questions."

Read those the way the model does. Four of them are three-word fragments carrying almost no signal. The fifth is a full sentence that ends with the phrase "answers to customer questions" — and the input is, unmistakably, a customer question. The best-described tool won, because tool selection is next-token prediction over the text you wrote, and the text you wrote said that one was for this.

This is the single most important thing to understand about tool use: a tool definition is not documentation that sits beside the model. It is prompt text inside the model's context, competing for attention with everything else. Debug tool selection the way you debug a prompt, because it is one.

Four calling patterns, and who sequences themyou1 callnonethe model0 or 1calls it out of habitthe model3 or morestep two on a stale idyou1 roundone failuresinks the batchWho decidesCallsFailure to watchDirectConditionalChainedParallelTool schemas are re-sent and re-billed on every request, so an unused tool costs money on each turn.
Sequence the calls yourself whenever the order is knowable — the model only needs to choose when you genuinely cannot.

What a tool definition actually is

When you pass a tools array, it is serialised and placed in the context ahead of the conversation. The model sees names, descriptions and schemas as text. Then, at each turn, it either predicts ordinary text or predicts a structured tool-use block naming one of them.

Three consequences follow immediately, and each has a cost attached.

Tool schemas are billed on every single request

A moderately detailed tool — name, two-sentence description, five typed parameters with per-field descriptions — runs about 120 tokens serialised. With 40 tools that is 4,800 tokens prepended to every request. In a 12-turn agent loop the schemas alone account for

12×4,800=57,600 input tokens12 \times 4{,}800 = 57{,}600 \text{ input tokens}

— roughly 0.29 dollars at 5 dollars per million, for one user question, before a single word of conversation. Prompt caching removes almost all of this, because the tool block is byte-identical on every turn. But only if it really is byte-identical: build the array in a deterministic order, with no timestamps, no dictionary iteration order, no per-request identifiers. One reordered key invalidates the cached prefix and you silently pay full price again.

Tool results are context, not return values

A tool result is pasted back as tokens. A search_orders call that returns 40 full order records at roughly 300 tokens each is 12,000 tokens injected into the conversation — which every subsequent turn must then re-send and attend past. Project the result in code first:

Python
def search_orders(customer_id, status=None, limit=20):    rows = db.query(...)    return [{"id": r.id, "date": str(r.date), "status": r.status,             "total": r.total, "items": len(r.items)} for r in rows[:limit]]

Five fields per order is about 20 tokens, so 40 orders becomes 800 tokens instead of 12,000 — a 15× reduction, with no loss of anything the model needed for the next decision. If it turns out to need the full record for one order, it can call lookup_order on that id. Return the minimum that supports the next decision, not everything the row contains.

The schema constrains the shape, never the meaning

With strict: true and additionalProperties: false, decoding is constrained so the emitted arguments are guaranteed to validate. That removes parse failures entirely — a genuine win, and you should switch it on. It does not guarantee the arguments are right. {"customer_id": "4471"} validates perfectly when 4471 was the order number.

Schema validation catches malformed calls. Only the description catches wrong ones. Effort spent on the description outperforms effort spent on the schema, every time.

Writing a definition the model can act on

Python
BAD = {    "name": "search",    "description": "Search for information.",    "input_schema": {        "type": "object",        "properties": {"query": {"type": "string"}},        "required": ["query"],    },}

The name is a generic verb, so it is a plausible answer to almost any need. The description states no domain, so nothing distinguishes it from any other search tool. There is no statement of what it returns, so the model cannot predict whether the result will help. And query as a bare string invites free-text where the backend may want an identifier.

Python
GOOD = {    "name": "search_orders",    "description": (        "Search a customer's order history. Returns id, date, status, total and "        "item count for up to `limit` orders, newest first.\n"        "USE WHEN: you need to find an order but do not know its id, or you need "        "several orders (e.g. 'my last three orders', 'anything still shipping').\n"        "DO NOT USE WHEN: you already have an order id — call `lookup_order` "        "instead, which returns line items, addresses and tracking.\n"        "Covers orders from 2021 onward only. Returns an empty list, not an error, "        "when nothing matches."    ),    "input_schema": {        "type": "object",        "properties": {            "customer_id": {                "type": "string",                "description": "Internal customer id, format CUS-123456. NOT an email "                               "address and NOT an order number. If you only have an "                               "email, call `find_customer` first.",            },            "status": {                "type": "string",                "enum": ["placed", "shipped", "delivered", "cancelled", "any"],                "description": "Filter by fulfilment status. Use 'any' for no filter.",            },            "limit": {"type": "integer", "minimum": 1, "maximum": 50, "default": 20},        },        "required": ["customer_id", "status"],        "additionalProperties": False,    },    "strict": True,}
ElementWhat it prevents
Specific verb_noun nameGeneric names attract calls from unrelated intents
"Returns id, date, status, total, item count"The model chaining onwards on fields that will not be there
USE WHEN with a concrete phrasingThe tool being skipped because its applicability was implicit
DO NOT USE WHEN, naming the siblingThe overlap failure that opened this lesson
"Covers 2021 onward"An empty result being read as "this customer has no orders"
"Returns an empty list, not an error"The model treating a valid zero-result as a failure to retry
enum instead of a free stringInvented values like "in transit" that the backend rejects
Field description ruling out an emailThe type-confusion 400 that costs a whole round trip
Pointer to find_customerThe model guessing an id it does not have

Every one of those lines exists because of a failure it prevents. That is the right way to grow a description: not by writing more, but by writing one sentence per observed mistake.

Four calling patterns

1. Direct call

One tool, one round, then answer. Use it when the tool is unambiguously required — a unit conversion, a lookup by a known key. There is nothing to decide, so do not build a loop.

2. Conditional use — deciding whether to call at all

This is where most agents leak errors, because the usual instruction is empty:

Text
BADUse the tools when you need them. If you can answer from your ownknowledge, do that instead.

"If you can answer from your own knowledge" fails because the model's estimate of whether it knows something is itself a generation, and a fluent completion is always available. It will produce a plausible refund window, a plausible price, a plausible policy. And the base rate is against you: the overwhelming majority of question-and-answer text in the training distribution contains no tool call at all, so the default pull is towards answering directly.

Text
GOODDecide by CATEGORY OF FACT, not by how confident you feel.MUST come from a tool — you do not know these and cannot infer them:  - anything about a specific customer, order, ticket or invoice  - current prices, stock levels, dates or balances  - any company policy, threshold or entitlement  - anything the user asserts about their account (verify it)MAY be answered directly:  - how a general feature works, described in the docs you were given  - definitions, explanations, arithmetic on numbers already in contextIf a needed fact falls in the first list and no tool provides it,say so plainly. Do not estimate it.

The difference is that the second version replaces an unanswerable question ("do I know this?") with a decidable one ("what kind of fact is this?"). Categorising a fact is a classification the model can actually do; introspecting on its own knowledge is not.

3. Chaining — and who should sequence it

When one tool's output feeds the next, you have a choice that people rarely make consciously.

Code-orchestratedModel-orchestrated
Who decides the orderYour programThe model, one step at a time
Model calls neededOne, at the endOne per step
Cost for a 4-step chain1 call4 calls, each resending the whole transcript
ReliabilityDeterministicVaries run to run
Handles branchingOnly branches you wroteAny branch it can reason to
Use whenThe sequence is known in advanceStep n+1n+1 genuinely depends on what step nn returned

The rule that follows: if you can draw the flowchart, write the flowchart. "Find the customer by email, then get their last order, then get its tracking" is a fixed sequence. Handing it to a model turns three deterministic function calls into three stochastic decisions, each with a failure probability, at four times the cost. Reserve model orchestration for the genuine case — where what you do next depends on what came back.

4. Parallel calls

A single turn may contain several independent tool-use blocks. Execute them concurrently, then return all results in one user message:

Python
results = []for block in (b for b in resp.content if b.type == "tool_use"):    try:        results.append({"type": "tool_result", "tool_use_id": block.id,                        "content": json.dumps(TOOLS[block.name](block.input))})    except Exception as e:        results.append({"type": "tool_result", "tool_use_id": block.id,                        "content": f"{type(e).__name__}: {e}", "is_error": True})messages.append({"role": "user", "content": results})     # one message, all results

Splitting those results across separate messages is accepted by the API and quietly teaches the model to stop emitting parallel calls — a slow performance regression with no error to point at. Dropping the result for a failed call is worse: it leaves a tool-use block with no matching result, which the model cannot reason about at all. Every emitted call gets a result, even if that result is an error.

Errors: which ones the model should ever see

The organising question is whether the model can do anything about it. If the fix is deterministic, fix it in code and never spend a round trip. If the fix requires a different decision, surface it — and phrase it as an instruction.

ClassExampleWhere handledWhat the model should see
TransientTimeout, 503, connection resetCode — retry with backoffNothing, unless retries are exhausted
Rate limited429Code — queue and waitNothing
Bad argumentscustomer_id was an emailModelThe exact validation message plus the correct format and the tool that produces it
Wrong toolOrder id passed to a customer lookupModelWhat was wrong and which tool to use instead
Zero resultsSearch matched nothingModelAn explicit statement that the query ran and matched nothing, with the coverage caveat
Permission denied403 on a restricted recordModelThat it is denied and will stay denied — do not retry
Malformed backend responseUpstream returned HTMLCode — normalise or fail cleanlyA clean typed error, never raw HTML

Write error text as an instruction

Text
BAD   {"error": "ValidationError: invalid input"}GOOD  "ValidationError: customer_id must match CUS-###### but got       'jane@example.com'. That looks like an email address. Call       find_customer(email='jane@example.com') to get the id, then       retry this call with it."

The second version is the next forward pass's context. It states the constraint, names the observed value, diagnoses the confusion, and gives the exact recovery call. Recovery from the first version requires the model to guess all four.

The empty-result trap

This one causes real incidents. A search returns nothing and you send back [] or null. The most probable continuation after an empty result is not "I could not find that" — it is the model filling the gap from its prior, because a helpful answer is far more likely in the training distribution than an admission of failure.

Text
BAD   []GOOD  {"results": [], "query_executed": "customer_id=CUS-004471,       status=any", "match_count": 0, "note": "The query ran       successfully and matched zero orders. This index covers       orders placed from 2021 onward. Zero matches means this       customer has no orders in that period — it does not mean       the lookup failed. Do not infer order details."}

You are not being verbose for its own sake. Each clause blocks a specific wrong continuation: that the tool broke, that older orders might exist and were missed, and that plausible details may be supplied.

Make writes idempotent

Retries happen — from your backoff logic, from a model that did not see the first result, from a user clicking twice. Any tool that mutates state needs a client-supplied key:

Python
def issue_refund(order_id: str, amount_pence: int, idempotency_key: str):    if (prior := refunds.get(idempotency_key)):        return {**prior, "note": "Already processed under this key; not repeated."}    ...

Then instruct the model to reuse the key when retrying the same logical action. Without this, "the tool timed out, let me try again" issues two refunds.

Tool discovery when there are too many tools

Selection accuracy degrades as the tool list grows — partly because descriptions start to overlap, partly because 40 tools is 4,800 tokens of context that the actual question has to compete with. Four approaches, in ascending order of complexity:

ApproachHowGood forCost
Flat listPass everythingUp to roughly 15 well-differentiated toolsAll schemas on every request
RoutingA cheap first call picks a category; a second call is issued with only that category's tools loadedTools that partition cleanly by domainOne extra small call; a routing mistake is unrecoverable
Deferred loading with a search toolMark most tools defer_loading: true and add a tool-search tool; the model retrieves definitions on demandLarge, sparse tool setsOnly retrieved schemas enter context; costs a search round trip
ConsolidationMerge near-duplicates behind one tool with an enum parameterFamilies like get_x_by_id repeated eight timesFree — usually the right first move

Deferred loading is worth knowing concretely. You mark the bulk of your tools with defer_loading: true so their schemas are not sent, and include a search tool — tool_search_tool_regex_20251119 or the BM25 variant — which the model uses to pull in the definitions it needs. Two constraints bite: the search tool itself must never be deferred, and at least one tool in the array must be non-deferred, or the request is rejected outright.

Before reaching for any of that, try consolidation. Eight tools named get_order_by_id, get_customer_by_id, get_ticket_by_id and so on are one tool with an entity enum. That removes seven descriptions the model must disambiguate between and roughly 800 tokens from every request, and it makes the remaining decision easier rather than harder.

A worked research agent

Four read-only tools, given deliberately non-overlapping scopes:

Python
TOOLS = [  {"name": "search_papers",   "description": "Full-text search over the internal paper index (abstracts and "                  "metadata only, not full text). Returns up to 10 hits with "                  "paper_id, title, year, venue, citation_count. Use to FIND papers. "                  "To read one, call fetch_paper with its paper_id."},  {"name": "fetch_paper",   "description": "Fetch one paper by paper_id. Returns abstract, methods section and "                  "reported results. Roughly 800 tokens. Never call this on more than "                  "three papers before summarising what you have."},  {"name": "run_python",   "description": "Execute a Python snippet in a sandbox with numpy and pandas; returns "                  "stdout. Use for ALL arithmetic, statistics and unit conversion. Do "                  "not compute multi-digit arithmetic yourself."},  {"name": "web_search",   "description": "Public web search for anything published after the internal index "                  "cut-off, or for material outside the index (news, standards, "                  "vendor documentation). Prefer search_papers for academic work."},]

Task: "Have reported error rates on the benchmark improved since 2022, and by how much?"

Text
Thought: I need reported numbers over time, so first find thepapers, then read only the ones that report a rate.Action: search_papers {"query": "benchmark error rate", "year_from": 2022}Observation: 7 hits. [{paper_id: P-1102, year: 2022, ...}, ...]Thought: 7 hits but I may only read three. The 2022 and 2025papers bracket the range; add the most-cited one in between.Action: fetch_paper {"paper_id": "P-1102"}Observation: {abstract: ..., results: "error rate 14.2% on the             standard split (n=4,000)"}Thought: 14.2% in 2022. Now the most recent.Action: fetch_paper {"paper_id": "P-2440"}Observation: {results: "error rate 6.8% on the standard split             (n=4,000)"}Thought: 14.2 to 6.8. I must not compute the relative change in myhead — that is a tool job.Action: run_python {"code": "print(round((14.2-6.8)/14.2*100, 1))"}Observation: 52.1Final Answer: Reported error rate on the standard split fell from14.2% (2022, P-1102) to 6.8% (2025, P-2440) — an absolute drop of7.4 points and a 52.1% relative reduction. Both figures are on thesame 4,000-example split. Note this compares two papers' self-reported numbers, not a controlled re-evaluation.

Three design choices produced that behaviour, and all three live in the tool descriptions. The fetch_paper budget ("never more than three") stopped the agent pulling 7 × 800 = 5,600 tokens of papers into context before doing anything with them. The run_python description ("do not compute multi-digit arithmetic yourself") moved a serial algorithm onto hardware that does serial algorithms exactly — digit-by-digit carrying is a poor fit for a fixed number of layers per token. And the explicit division of labour between search_papers and web_search stopped the two competing for the same intent.

Designing the tool surface itself

The deepest decision is granularity, and it is a genuine trade-off rather than a best practice.

Broad tools (run_sql, bash)Narrow tools (get_overdue_invoices)
Context costTiny — one schemaGrows linearly with the tool count
CoverageAnything expressible in the languageOnly what you anticipated
Failure modeSyntactically valid, semantically wrong — a join on the wrong key, a DELETE without a WHEREThe needed operation simply does not exist
AuditabilityPoor — the intent must be inferred from generated codeExcellent — the tool name is the intent
SafetyRequires sandboxing, permissions, query reviewBounded by construction
Best forExploratory work by expert users, read-only analysisProduction paths, anything that writes

A useful hybrid: narrow tools for every write and every high-frequency read, plus one broad read-only escape hatch for the long tail — running against a replica, with a row cap and a statement timeout. You get bounded, auditable behaviour on the paths that matter and coverage everywhere else.

Schema features that earn their place

  • enum for any closed set. It is the difference between a validated choice and a guessed string, and it also documents the options in the same breath.
  • minimum / maximum on numbers. A limit with no maximum will eventually be called with 10000.
  • additionalProperties: false with strict: true. Guarantees the emitted call validates, so the argument-shape error class disappears rather than being handled.
  • A description on every field, not just on the tool. Field descriptions are read at exactly the moment the value is being generated.
  • Flat objects. Deep nesting means more structural tokens to get right and more ways to get them subtly wrong.
  • Parse tool inputs with json.loads, never with string matching — escaping of Unicode and forward slashes varies between models and a raw substring check will pass in testing and fail in production.

What this means when you build something

Two habits separate tool integrations that work from ones that mostly work.

Instrument selection, then fix the descriptions. Log every call: which tool, which arguments, whether it errored, whether the turn made progress. Then compute the numbers that matter — how often each tool is chosen, which pairs get confused, which tools are never chosen at all. A tool chosen 40% of the time and useful 10% of the time has an over-broad description. A tool never chosen is invisible, and the fix is one USE WHEN line, not a different model. This loop — measure confusion, add one sentence, measure again — is the highest-yield work available in agent development, and it is cheap.

Write every tool twice: once for the happy path, once for the model. The second version is the part people skip, and it is where the reliability lives. For each tool, ask: what does it return when there is nothing to return, and will that read as failure? What does its error message tell the model to do? What happens if it is called twice with the same arguments? Is what it returns proportionate to what the next decision needs, or is it 12,000 tokens of context that everything downstream has to read past?

Get those four answers right and most agent failures disappear before the loop is ever written. The support agent that opened this lesson needed one clause — "DO NOT USE for questions about a specific order; use lookup_order" — on a single description. Not a better model, not a bigger prompt. One sentence, in the place the model was actually reading.