Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Using APIs and Functions as Tools
An internal support agent had 23 tools registered. Someone asked it: "How many orders did customer 8842 place last month?"
It called search_documents("customer 8842 orders"). That tool searches the help-centre knowledge base — articles about how to place an order, how to cancel one, what the refund window is. It returned three articles. The agent read them, found no number, and answered: "I could not find order history for that customer."
The right tool was run_query, which runs read-only SQL against the orders database and would have answered in 40 milliseconds. Why did the agent not pick it?
search_documents: "Search for information."run_query: "Query the database."Read those two descriptions as the model sees them — with no knowledge of your architecture, no idea that "documents" means help articles, no idea what is in "the database". The question contains the word "orders". One tool mentions searching for information, which is what the user wants. The model picked reasonably from what it was given.
The bug was not in the model. It was in two lines of text a developer wrote in four seconds.
What a tool actually is
A tool is a function the model may ask you to call. That is the entire mechanism, and it is worth being precise about who does what:
1. You send: the conversation + a list of tool SCHEMAS2. Model returns: "call get_order(order_id='ORD-4471')"3. YOUR CODE: actually calls get_order('ORD-4471')4. You send: the conversation + the return value5. Model returns: an answer, or another tool callThe model never executes anything. It emits a structured request, and your code decides whether to honour it. Everything about tool design follows from that division: the model's only information about a tool is its schema, and your code's only protection against a bad request is validation.
The model does not see your function. It sees your description of it. If those two disagree, the description wins — and the failure appears in your code, not in the prompt.
Anatomy of a tool definition
Here is the same tool written badly and then properly.
Bad
1{2 "name": "get_data",3 "description": "Gets data",4 "parameters": {5 "type": "object",6 "properties": {7 "query": { "type": "string" },8 "opts": { "type": "object" }9 }10 }11}Four separate failures. The name says nothing. The description says less. query could be SQL, a search phrase, a customer ID, or a question in English — the model will guess, and different calls will guess differently. And opts as a free-form object is an invitation to invent keys that your code has never heard of.
Good
1{2 "name": "get_customer_orders",3 "description": "Return the order history for one customer from the orders database. Use this for questions about what a specific customer bought, when, and for how much. Does NOT search help articles or product descriptions - use search_help_articles for those. Returns at most 100 orders, newest first.",4 "parameters": {5 "type": "object",6 "properties": {7 "customer_id": {8 "type": "integer",9 "description": "Numeric internal customer ID, e.g. 8842. Not an email address and not an order ID."10 },11 "start_date": {12 "type": "string",13 "description": "Inclusive start of the range, ISO format YYYY-MM-DD, e.g. 2026-07-01."14 },15 "end_date": {16 "type": "string",17 "description": "Inclusive end of the range, ISO format YYYY-MM-DD."18 },19 "status": {20 "type": "string",21 "enum": ["placed", "shipped", "delivered", "cancelled", "any"],22 "description": "Filter by order status. Use 'any' for all statuses."23 }24 },25 "required": ["customer_id", "start_date", "end_date"]26 }27}Every addition is load-bearing:
| Element | Prevents |
|---|---|
get_customer_orders not get_data | Guessing what the tool is for from the description alone |
| "Use this for questions about what a specific customer bought" | Not recognising this tool applies to the question |
| "Does NOT search help articles — use search_help_articles" | The exact confusion that produced the failure above |
| "Returns at most 100 orders" | Concluding a customer has exactly 100 orders |
customer_id is integer | Passing "customer 8842" as a string |
| "Not an email address and not an order ID" | Passing "ORD-4471" into a customer field |
| ISO format spelled out with an example | "last month", "07/2026", "July" |
enum on status | Inventing "in transit" or "complete" |
required list | Omitting the date range and pulling all history |
The negative statement — "does NOT do X, use Y for that" — is the single highest-value sentence you can add when two tools are adjacent in meaning. It is worth writing it on both tools, in both directions.
Constraints the schema cannot express
JSON Schema can say "integer" and "one of these five values". It cannot say "end_date must be after start_date", "the range must be under 366 days", or "this customer must belong to the caller's organisation". Those go in two places, and you need both.
In the description, so the model can get it right first time:
"The range must be at most 366 days and end_date must not be before start_date. Requests outside these limits are rejected."In the code, so a wrong call fails safely and informatively:
1from datetime import date23MAX_RANGE_DAYS = 36645def get_customer_orders(customer_id, start_date, end_date, status="any"):6 try:7 start = date.fromisoformat(start_date)8 end = date.fromisoformat(end_date)9 except ValueError:10 return {"error": "invalid_date_format",11 "message": "Dates must be YYYY-MM-DD, e.g. 2026-07-01.",12 "received": {"start_date": start_date,13 "end_date": end_date}}1415 if end < start:16 return {"error": "invalid_range",17 "message": (f"end_date {end_date} is before start_date "18 f"{start_date}. Swap them.")}1920 span = (end - start).days21 if span > MAX_RANGE_DAYS:22 return {"error": "range_too_large",23 "message": (f"Requested {span} days; maximum is "24 f"{MAX_RANGE_DAYS}. Split into "25 f"{-(-span // MAX_RANGE_DAYS)} calls.")}2627 rows = db.fetch_orders(customer_id, start, end, status)28 return {"customer_id": customer_id,29 "order_count": len(rows),30 "truncated": len(rows) == 100,31 "orders": rows}Two things about those error returns. They are returned, not raised, so they arrive in the agent's observation channel where the model can read them. And each one says what to do next — swap the dates, split into two calls. An error message that only reports what went wrong makes the agent guess; one that names the fix makes it succeed on the retry.
The truncated flag matters just as much. Without it, an agent that receives exactly 100 orders will confidently report "this customer placed 100 orders". With it, the agent knows the answer is a floor, not a count.
The function-calling round trip
1import json, os2from anthropic import Anthropic34client = Anthropic()5MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5") # current default; set per deployment67# Anthropic's API calls the schema field input_schema rather than parameters.8TOOLS = [{"name": s["name"], "description": s["description"],9 "input_schema": s["parameters"]}10 for s in (GET_CUSTOMER_ORDERS_SCHEMA, SEARCH_HELP_ARTICLES_SCHEMA)]11REGISTRY = {"get_customer_orders": get_customer_orders,12 "search_help_articles": search_help_articles}1314messages = [{"role": "user",15 "content": "How many orders did customer 8842 place "16 "in July 2026?"}]1718while True:19 resp = client.messages.create(20 model=MODEL,21 max_tokens=1024,22 tools=TOOLS,23 messages=messages,24 )25 messages.append({"role": "assistant", "content": resp.content})2627 if resp.stop_reason != "tool_use":28 print("".join(b.text for b in resp.content if b.type == "text"))29 break3031 results = []32 for block in resp.content:33 if block.type != "tool_use":34 continue35 fn = REGISTRY.get(block.name)36 if fn is None:37 out = {"error": "unknown_tool",38 "message": f"No tool named {block.name}. "39 f"Available: {list(REGISTRY)}"}40 else:41 try:42 out = fn(**block.input)43 except TypeError as e:44 out = {"error": "bad_arguments", "message": str(e)}45 except Exception as e:46 out = {"error": type(e).__name__, "message": str(e)}47 results.append({"type": "tool_result",48 "tool_use_id": block.id,49 "content": json.dumps(out)})5051 messages.append({"role": "user", "content": results})Four details that are easy to get wrong:
- The assistant turn is appended before results. The tool-use block must remain in the conversation or the model loses track of what it asked for.
tool_use_idmust match. When a model requests three tools in one turn, this is what pairs each result with its request.- All results go back in one user message. Not three separate turns.
- Every failure path returns a dict. Unknown tool, wrong arguments, tool crash — all become observations, never exceptions that kill the loop.
The schemas in this lesson use the provider-neutral parameters key. Each provider wraps the same JSON Schema slightly differently — Anthropic's Messages API wants it as input_schema, OpenAI's function tools as parameters — which is why the code converts once, at the edge, instead of scattering provider formats through the tool library.
Organising a tool library
Past about eight tools, an unstructured list stops working. Three problems arrive together: the model picks wrong more often, the schemas eat your context window, and nobody on the team can remember what exists.
The token arithmetic
A well-written schema costs roughly 150–200 tokens. With 23 tools that is about 4,200 tokens re-sent on every model call. Over a 10-step agent run:
| Tools in prompt | Schema tokens/call | Over 10 steps | Cost at an example 3 USD per million input |
|---|---|---|---|
| 5 | ~900 | 9,000 | 2.7 cents |
| 12 | ~2,200 | 22,000 | 6.6 cents |
| 23 | ~4,200 | 42,000 | 12.6 cents |
| 60 | ~11,000 | 110,000 | 33 cents |
Thirty-three cents per run of pure overhead, before the agent has done anything, and multiplied across thousands of runs a day. Prompt caching reduces the price but not the context consumed.
The selection problem
The accuracy cost is worse than the token cost. Tool selection is a discrimination task: the model must pick one option from N descriptions. Every additional tool adds another chance for a near-miss, and near-misses cluster — a library with search_documents, search_articles, find_content and lookup_kb is four tools that are mutually confusable no matter how good the model is.
Three fixes, in order of how much they help:
Merge overlapping tools. Those four search tools are one tool with a source enum. One schema, one decision, no ambiguity:
1{2 "name": "search",3 "description": "Full-text search over one content source.",4 "parameters": {5 "type": "object",6 "properties": {7 "source": {8 "type": "string",9 "enum": ["help_articles", "product_catalogue",10 "internal_wiki", "past_tickets"],11 "description": "help_articles: customer-facing how-to guides. product_catalogue: SKUs, specs, prices. internal_wiki: engineering and policy docs. past_tickets: resolved support conversations."12 },13 "query": { "type": "string", "description": "Search phrase." }14 },15 "required": ["source", "query"]16 }17}Namespace by domain. orders_get, orders_cancel, billing_refund, billing_invoice. The prefix carries information before the model reads a single word of description, and it makes the library legible to humans too.
Load tools dynamically. Keep a compact index of all tools, embed the descriptions, and at each step retrieve only the 5–8 most relevant to the current goal:
1class ToolLibrary:2 def __init__(self, tools, embedder):3 self.tools = {t["name"]: t for t in tools}4 self.vectors = {t["name"]: embedder(t["description"])5 for t in tools}6 self.embedder = embedder78 def relevant(self, goal, k=6, always=("finish",)):9 q = self.embedder(goal)10 ranked = sorted(self.tools,11 key=lambda n: -cosine(q, self.vectors[n]))12 chosen = list(always) + [n for n in ranked if n not in always][:k]13 return [self.tools[n] for n in chosen]The always parameter is not decoration. Tools the agent needs regardless of topic — finishing, escalating, asking the user — must never be retrieved away, or the agent will occasionally be unable to stop.
Composing tools
Most real work chains tools, where one output becomes another's input. This only works if the types line up.
get_customer_orders(8842, ...) -> { orders: [ {order_id: "ORD-4471", total_paise: 425000, status: "delivered"} ] } │ ▼ order_idget_shipment(order_id="ORD-4471") -> { carrier: "Delhivery", tracking: "1Z994A", delivered_at: "2026-07-19" } │ ▼ trackingget_tracking_events(tracking="1Z994A")The chain holds because get_customer_orders returns order_id in exactly the format get_shipment expects. Break that and the agent must transform values itself — which it will do by guessing, and guessing about identifier formats fails silently.
| Composition problem | Symptom | Fix |
|---|---|---|
Tool A returns id: 4471, tool B wants "ORD-4471" | 404 on every chained call | Return the canonical form both tools use |
| A returns cents, B expects rupees | Amounts off by 100× | Put the unit in the field name everywhere |
| A returns a list, B takes one item | Agent loops one-by-one, burning steps | Give B a batch parameter |
| A returns 40KB of JSON | Context fills; the useful field is buried | Add a fields parameter; return summaries by default |
A's date is 2026-07-19, B wants epoch seconds | Silently wrong time windows | ISO 8601 strings at every boundary |
Where a chain is fixed and frequent, collapse it into one tool. If "customer ID to delivery status" is asked a hundred times a day, write get_delivery_status(customer_id) that runs all three internally. You turn three model calls into one, remove two chances to pick the wrong tool, and cut the token cost roughly threefold. The agent's judgement should be spent on decisions that are genuinely open, not on re-deriving a fixed sequence.
Every step you can move from the agent's reasoning into a tool's implementation is a step that becomes fast, cheap, testable and deterministic.
Real tool groupings
| Agent | Read tools | Write tools | Boundary to enforce |
|---|---|---|---|
| Support | orders_get, billing_invoices, search(source) | billing_refund, tickets_escalate, email_send | Refunds capped and logged; email templates fixed, not free text |
| Research | web_search, page_fetch, pdf_extract | notes_append | Domain allowlist; fetch size cap; no outbound POSTs |
| Data analysis | sql_select, schema_describe, chart_render | — none — | Read-only DB role; statement timeout; row cap |
| Coding | file_read, grep, tests_run | file_write, git_commit | Writes confined to the repo; no push; no network in tests |
| Ops | metrics_query, logs_search, service_status | service_restart, scale_set | Blast-radius limit; approval for anything customer-facing |
The pattern across every row: read tools far outnumber write tools, and every write tool has a named constraint. That asymmetry is deliberate. Reads are cheap to get wrong and writes are not, so give the agent generous read access and grudging write access.
Where people get this wrong
Writing the description for yourself. "Gets order data" is meaningful to you because you know what an order is here. Write for a competent stranger who has your schema and nothing else. If two people on your team would interpret the description differently, so will the model.
Exposing raw API responses. A payment provider returning 60 fields of which 3 matter costs you tokens on every call and hides the signal. Wrap it, return what the agent needs, and add a fields parameter for the rare case where it needs more.
Raising exceptions instead of returning errors. An exception ends the loop. A returned error dict keeps the agent alive and gives it something to act on. The only things that should raise are conditions where continuing would be unsafe.
One tool that does everything. The mirror image of too many tools: a single execute(command) that takes a natural-language instruction. It looks elegant and it is untestable, unvalidatable, and impossible to permission — you cannot grant read-only access to a tool whose scope is "anything".
Trusting arguments because a model produced them. Model output is untrusted input, and it can be influenced by text in the conversation. Validate types, ranges, ownership and permission on every call, exactly as you would for a request arriving from the internet — because in effect it is one.
Silent truncation. A tool that caps results at 100 and does not say so turns a limit into a lie. Any cap, filter or approximation must appear in the return value.
What this means when you build one
Write the tool schema before the implementation. Naming the tool, describing when to use it and when not to, and enumerating the parameter values forces you to decide what the tool is for — and that decision is the one the model inherits.
Then run the cheapest useful test there is: put your tool list in front of a colleague who does not know the system, give them ten real user questions, and ask which tool they would call. Every disagreement is a description bug, and you have found it for the price of ten minutes rather than a week of confused production traces.
Keep the active tool list under about eight. Merge near-duplicates behind an enum, namespace what remains by domain, and retrieve dynamically if the full library is large — pinning the tools that must always be available.
Make every return value structured, unit-suffixed, and honest about its own limits: truncated, as_of, excluded, approximate. And write your error returns as instructions. The difference between {"error": "invalid_range"} and {"error": "invalid_range", "message": "end_date 2026-07-01 is before start_date 2026-07-31. Swap them."} is the difference between an agent that fails and an agent that fixes itself on the next turn — which is, in the end, what you were building the loop for.