Course Content
Advanced Prompting and Reasoning
3 sections · 7 lessons
Tool-Augmented Reasoning
A support agent has five tools. A customer types: "Where is my order #4471?"
The agent calls search_knowledge_base("where is my order") and returns a help article titled "Understanding Shipping Times". The order status was one lookup_order call away and the agent never made it.
Nobody wrote a bug. The tool descriptions looked like this:
lookup_order "Look up an order."search_orders "Search orders."search_customers "Search customers."search_tickets "Search support tickets."search_knowledge_base "Search the knowledge base for answers to customer questions."Read those the way the model does. Four of them are three-word fragments carrying almost no signal. The fifth is a full sentence that ends with the phrase "answers to customer questions" — and the input is, unmistakably, a customer question. The best-described tool won, because tool selection is next-token prediction over the text you wrote, and the text you wrote said that one was for this.
This is the single most important thing to understand about tool use: a tool definition is not documentation that sits beside the model. It is prompt text inside the model's context, competing for attention with everything else. Debug tool selection the way you debug a prompt, because it is one.
What a tool definition actually is
When you pass a tools array, it is serialised and placed in the context ahead of the conversation. The model sees names, descriptions and schemas as text. Then, at each turn, it either predicts ordinary text or predicts a structured tool-use block naming one of them.
Three consequences follow immediately, and each has a cost attached.
Tool schemas are billed on every single request
A moderately detailed tool — name, two-sentence description, five typed parameters with per-field descriptions — runs about 120 tokens serialised. With 40 tools that is 4,800 tokens prepended to every request. In a 12-turn agent loop the schemas alone account for
— roughly 0.29 dollars at 5 dollars per million, for one user question, before a single word of conversation. Prompt caching removes almost all of this, because the tool block is byte-identical on every turn. But only if it really is byte-identical: build the array in a deterministic order, with no timestamps, no dictionary iteration order, no per-request identifiers. One reordered key invalidates the cached prefix and you silently pay full price again.
Tool results are context, not return values
A tool result is pasted back as tokens. A search_orders call that returns 40 full order records at roughly 300 tokens each is 12,000 tokens injected into the conversation — which every subsequent turn must then re-send and attend past. Project the result in code first:
1def search_orders(customer_id, status=None, limit=20):2 rows = db.query(...)3 return [{"id": r.id, "date": str(r.date), "status": r.status,4 "total": r.total, "items": len(r.items)} for r in rows[:limit]]Five fields per order is about 20 tokens, so 40 orders becomes 800 tokens instead of 12,000 — a 15× reduction, with no loss of anything the model needed for the next decision. If it turns out to need the full record for one order, it can call lookup_order on that id. Return the minimum that supports the next decision, not everything the row contains.
The schema constrains the shape, never the meaning
With strict: true and additionalProperties: false, decoding is constrained so the emitted arguments are guaranteed to validate. That removes parse failures entirely — a genuine win, and you should switch it on. It does not guarantee the arguments are right. {"customer_id": "4471"} validates perfectly when 4471 was the order number.
Schema validation catches malformed calls. Only the description catches wrong ones. Effort spent on the description outperforms effort spent on the schema, every time.
Writing a definition the model can act on
1BAD = {2 "name": "search",3 "description": "Search for information.",4 "input_schema": {5 "type": "object",6 "properties": {"query": {"type": "string"}},7 "required": ["query"],8 },9}The name is a generic verb, so it is a plausible answer to almost any need. The description states no domain, so nothing distinguishes it from any other search tool. There is no statement of what it returns, so the model cannot predict whether the result will help. And query as a bare string invites free-text where the backend may want an identifier.
1GOOD = {2 "name": "search_orders",3 "description": (4 "Search a customer's order history. Returns id, date, status, total and "5 "item count for up to `limit` orders, newest first.\n"6 "USE WHEN: you need to find an order but do not know its id, or you need "7 "several orders (e.g. 'my last three orders', 'anything still shipping').\n"8 "DO NOT USE WHEN: you already have an order id — call `lookup_order` "9 "instead, which returns line items, addresses and tracking.\n"10 "Covers orders from 2021 onward only. Returns an empty list, not an error, "11 "when nothing matches."12 ),13 "input_schema": {14 "type": "object",15 "properties": {16 "customer_id": {17 "type": "string",18 "description": "Internal customer id, format CUS-123456. NOT an email "19 "address and NOT an order number. If you only have an "20 "email, call `find_customer` first.",21 },22 "status": {23 "type": "string",24 "enum": ["placed", "shipped", "delivered", "cancelled", "any"],25 "description": "Filter by fulfilment status. Use 'any' for no filter.",26 },27 "limit": {"type": "integer", "minimum": 1, "maximum": 50, "default": 20},28 },29 "required": ["customer_id", "status"],30 "additionalProperties": False,31 },32 "strict": True,33}| Element | What it prevents |
|---|---|
Specific verb_noun name | Generic names attract calls from unrelated intents |
| "Returns id, date, status, total, item count" | The model chaining onwards on fields that will not be there |
| USE WHEN with a concrete phrasing | The tool being skipped because its applicability was implicit |
| DO NOT USE WHEN, naming the sibling | The overlap failure that opened this lesson |
| "Covers 2021 onward" | An empty result being read as "this customer has no orders" |
| "Returns an empty list, not an error" | The model treating a valid zero-result as a failure to retry |
enum instead of a free string | Invented values like "in transit" that the backend rejects |
| Field description ruling out an email | The type-confusion 400 that costs a whole round trip |
Pointer to find_customer | The model guessing an id it does not have |
Every one of those lines exists because of a failure it prevents. That is the right way to grow a description: not by writing more, but by writing one sentence per observed mistake.
Four calling patterns
1. Direct call
One tool, one round, then answer. Use it when the tool is unambiguously required — a unit conversion, a lookup by a known key. There is nothing to decide, so do not build a loop.
2. Conditional use — deciding whether to call at all
This is where most agents leak errors, because the usual instruction is empty:
BADUse the tools when you need them. If you can answer from your ownknowledge, do that instead."If you can answer from your own knowledge" fails because the model's estimate of whether it knows something is itself a generation, and a fluent completion is always available. It will produce a plausible refund window, a plausible price, a plausible policy. And the base rate is against you: the overwhelming majority of question-and-answer text in the training distribution contains no tool call at all, so the default pull is towards answering directly.
GOODDecide by CATEGORY OF FACT, not by how confident you feel.MUST come from a tool — you do not know these and cannot infer them: - anything about a specific customer, order, ticket or invoice - current prices, stock levels, dates or balances - any company policy, threshold or entitlement - anything the user asserts about their account (verify it)MAY be answered directly: - how a general feature works, described in the docs you were given - definitions, explanations, arithmetic on numbers already in contextIf a needed fact falls in the first list and no tool provides it,say so plainly. Do not estimate it.The difference is that the second version replaces an unanswerable question ("do I know this?") with a decidable one ("what kind of fact is this?"). Categorising a fact is a classification the model can actually do; introspecting on its own knowledge is not.
3. Chaining — and who should sequence it
When one tool's output feeds the next, you have a choice that people rarely make consciously.
| Code-orchestrated | Model-orchestrated | |
|---|---|---|
| Who decides the order | Your program | The model, one step at a time |
| Model calls needed | One, at the end | One per step |
| Cost for a 4-step chain | 1 call | 4 calls, each resending the whole transcript |
| Reliability | Deterministic | Varies run to run |
| Handles branching | Only branches you wrote | Any branch it can reason to |
| Use when | The sequence is known in advance | Step n+1 genuinely depends on what step n returned |
The rule that follows: if you can draw the flowchart, write the flowchart. "Find the customer by email, then get their last order, then get its tracking" is a fixed sequence. Handing it to a model turns three deterministic function calls into three stochastic decisions, each with a failure probability, at four times the cost. Reserve model orchestration for the genuine case — where what you do next depends on what came back.
4. Parallel calls
A single turn may contain several independent tool-use blocks. Execute them concurrently, then return all results in one user message:
1results = []2for block in (b for b in resp.content if b.type == "tool_use"):3 try:4 results.append({"type": "tool_result", "tool_use_id": block.id,5 "content": json.dumps(TOOLS[block.name](block.input))})6 except Exception as e:7 results.append({"type": "tool_result", "tool_use_id": block.id,8 "content": f"{type(e).__name__}: {e}", "is_error": True})9messages.append({"role": "user", "content": results}) # one message, all resultsSplitting those results across separate messages is accepted by the API and quietly teaches the model to stop emitting parallel calls — a slow performance regression with no error to point at. Dropping the result for a failed call is worse: it leaves a tool-use block with no matching result, which the model cannot reason about at all. Every emitted call gets a result, even if that result is an error.
Errors: which ones the model should ever see
The organising question is whether the model can do anything about it. If the fix is deterministic, fix it in code and never spend a round trip. If the fix requires a different decision, surface it — and phrase it as an instruction.
| Class | Example | Where handled | What the model should see |
|---|---|---|---|
| Transient | Timeout, 503, connection reset | Code — retry with backoff | Nothing, unless retries are exhausted |
| Rate limited | 429 | Code — queue and wait | Nothing |
| Bad arguments | customer_id was an email | Model | The exact validation message plus the correct format and the tool that produces it |
| Wrong tool | Order id passed to a customer lookup | Model | What was wrong and which tool to use instead |
| Zero results | Search matched nothing | Model | An explicit statement that the query ran and matched nothing, with the coverage caveat |
| Permission denied | 403 on a restricted record | Model | That it is denied and will stay denied — do not retry |
| Malformed backend response | Upstream returned HTML | Code — normalise or fail cleanly | A clean typed error, never raw HTML |
Write error text as an instruction
BAD {"error": "ValidationError: invalid input"}GOOD "ValidationError: customer_id must match CUS-###### but got 'jane@example.com'. That looks like an email address. Call find_customer(email='jane@example.com') to get the id, then retry this call with it."The second version is the next forward pass's context. It states the constraint, names the observed value, diagnoses the confusion, and gives the exact recovery call. Recovery from the first version requires the model to guess all four.
The empty-result trap
This one causes real incidents. A search returns nothing and you send back [] or null. The most probable continuation after an empty result is not "I could not find that" — it is the model filling the gap from its prior, because a helpful answer is far more likely in the training distribution than an admission of failure.
BAD []GOOD {"results": [], "query_executed": "customer_id=CUS-004471, status=any", "match_count": 0, "note": "The query ran successfully and matched zero orders. This index covers orders placed from 2021 onward. Zero matches means this customer has no orders in that period — it does not mean the lookup failed. Do not infer order details."}You are not being verbose for its own sake. Each clause blocks a specific wrong continuation: that the tool broke, that older orders might exist and were missed, and that plausible details may be supplied.
Make writes idempotent
Retries happen — from your backoff logic, from a model that did not see the first result, from a user clicking twice. Any tool that mutates state needs a client-supplied key:
1def issue_refund(order_id: str, amount_pence: int, idempotency_key: str):2 if (prior := refunds.get(idempotency_key)):3 return {**prior, "note": "Already processed under this key; not repeated."}4 ...Then instruct the model to reuse the key when retrying the same logical action. Without this, "the tool timed out, let me try again" issues two refunds.
Tool discovery when there are too many tools
Selection accuracy degrades as the tool list grows — partly because descriptions start to overlap, partly because 40 tools is 4,800 tokens of context that the actual question has to compete with. Four approaches, in ascending order of complexity:
| Approach | How | Good for | Cost |
|---|---|---|---|
| Flat list | Pass everything | Up to roughly 15 well-differentiated tools | All schemas on every request |
| Routing | A cheap first call picks a category; a second call is issued with only that category's tools loaded | Tools that partition cleanly by domain | One extra small call; a routing mistake is unrecoverable |
| Deferred loading with a search tool | Mark most tools defer_loading: true and add a tool-search tool; the model retrieves definitions on demand | Large, sparse tool sets | Only retrieved schemas enter context; costs a search round trip |
| Consolidation | Merge near-duplicates behind one tool with an enum parameter | Families like get_x_by_id repeated eight times | Free — usually the right first move |
Deferred loading is worth knowing concretely. You mark the bulk of your tools with defer_loading: true so their schemas are not sent, and include a search tool — tool_search_tool_regex_20251119 or the BM25 variant — which the model uses to pull in the definitions it needs. Two constraints bite: the search tool itself must never be deferred, and at least one tool in the array must be non-deferred, or the request is rejected outright.
Before reaching for any of that, try consolidation. Eight tools named get_order_by_id, get_customer_by_id, get_ticket_by_id and so on are one tool with an entity enum. That removes seven descriptions the model must disambiguate between and roughly 800 tokens from every request, and it makes the remaining decision easier rather than harder.
A worked research agent
Four read-only tools, given deliberately non-overlapping scopes:
1TOOLS = [2 {"name": "search_papers",3 "description": "Full-text search over the internal paper index (abstracts and "4 "metadata only, not full text). Returns up to 10 hits with "5 "paper_id, title, year, venue, citation_count. Use to FIND papers. "6 "To read one, call fetch_paper with its paper_id."},7 {"name": "fetch_paper",8 "description": "Fetch one paper by paper_id. Returns abstract, methods section and "9 "reported results. Roughly 800 tokens. Never call this on more than "10 "three papers before summarising what you have."},11 {"name": "run_python",12 "description": "Execute a Python snippet in a sandbox with numpy and pandas; returns "13 "stdout. Use for ALL arithmetic, statistics and unit conversion. Do "14 "not compute multi-digit arithmetic yourself."},15 {"name": "web_search",16 "description": "Public web search for anything published after the internal index "17 "cut-off, or for material outside the index (news, standards, "18 "vendor documentation). Prefer search_papers for academic work."},19]Task: "Have reported error rates on the benchmark improved since 2022, and by how much?"
Thought: I need reported numbers over time, so first find thepapers, then read only the ones that report a rate.Action: search_papers {"query": "benchmark error rate", "year_from": 2022}Observation: 7 hits. [{paper_id: P-1102, year: 2022, ...}, ...]Thought: 7 hits but I may only read three. The 2022 and 2025papers bracket the range; add the most-cited one in between.Action: fetch_paper {"paper_id": "P-1102"}Observation: {abstract: ..., results: "error rate 14.2% on the standard split (n=4,000)"}Thought: 14.2% in 2022. Now the most recent.Action: fetch_paper {"paper_id": "P-2440"}Observation: {results: "error rate 6.8% on the standard split (n=4,000)"}Thought: 14.2 to 6.8. I must not compute the relative change in myhead — that is a tool job.Action: run_python {"code": "print(round((14.2-6.8)/14.2*100, 1))"}Observation: 52.1Final Answer: Reported error rate on the standard split fell from14.2% (2022, P-1102) to 6.8% (2025, P-2440) — an absolute drop of7.4 points and a 52.1% relative reduction. Both figures are on thesame 4,000-example split. Note this compares two papers' self-reported numbers, not a controlled re-evaluation.Three design choices produced that behaviour, and all three live in the tool descriptions. The fetch_paper budget ("never more than three") stopped the agent pulling 7 × 800 = 5,600 tokens of papers into context before doing anything with them. The run_python description ("do not compute multi-digit arithmetic yourself") moved a serial algorithm onto hardware that does serial algorithms exactly — digit-by-digit carrying is a poor fit for a fixed number of layers per token. And the explicit division of labour between search_papers and web_search stopped the two competing for the same intent.
Designing the tool surface itself
The deepest decision is granularity, and it is a genuine trade-off rather than a best practice.
Broad tools (run_sql, bash) | Narrow tools (get_overdue_invoices) | |
|---|---|---|
| Context cost | Tiny — one schema | Grows linearly with the tool count |
| Coverage | Anything expressible in the language | Only what you anticipated |
| Failure mode | Syntactically valid, semantically wrong — a join on the wrong key, a DELETE without a WHERE | The needed operation simply does not exist |
| Auditability | Poor — the intent must be inferred from generated code | Excellent — the tool name is the intent |
| Safety | Requires sandboxing, permissions, query review | Bounded by construction |
| Best for | Exploratory work by expert users, read-only analysis | Production paths, anything that writes |
A useful hybrid: narrow tools for every write and every high-frequency read, plus one broad read-only escape hatch for the long tail — running against a replica, with a row cap and a statement timeout. You get bounded, auditable behaviour on the paths that matter and coverage everywhere else.
Schema features that earn their place
enumfor any closed set. It is the difference between a validated choice and a guessed string, and it also documents the options in the same breath.minimum/maximumon numbers. Alimitwith no maximum will eventually be called with 10000.additionalProperties: falsewithstrict: true. Guarantees the emitted call validates, so the argument-shape error class disappears rather than being handled.- A
descriptionon every field, not just on the tool. Field descriptions are read at exactly the moment the value is being generated. - Flat objects. Deep nesting means more structural tokens to get right and more ways to get them subtly wrong.
- Parse tool inputs with
json.loads, never with string matching — escaping of Unicode and forward slashes varies between models and a raw substring check will pass in testing and fail in production.
What this means when you build something
Two habits separate tool integrations that work from ones that mostly work.
Instrument selection, then fix the descriptions. Log every call: which tool, which arguments, whether it errored, whether the turn made progress. Then compute the numbers that matter — how often each tool is chosen, which pairs get confused, which tools are never chosen at all. A tool chosen 40% of the time and useful 10% of the time has an over-broad description. A tool never chosen is invisible, and the fix is one USE WHEN line, not a different model. This loop — measure confusion, add one sentence, measure again — is the highest-yield work available in agent development, and it is cheap.
Write every tool twice: once for the happy path, once for the model. The second version is the part people skip, and it is where the reliability lives. For each tool, ask: what does it return when there is nothing to return, and will that read as failure? What does its error message tell the model to do? What happens if it is called twice with the same arguments? Is what it returns proportionate to what the next decision needs, or is it 12,000 tokens of context that everything downstream has to read past?
Get those four answers right and most agent failures disappear before the loop is ever written. The support agent that opened this lesson needed one clause — "DO NOT USE for questions about a specific order; use lookup_order" — on a single description. Not a better model, not a bigger prompt. One sentence, in the place the model was actually reading.