AI Agent Fundamentals

Using APIs and Functions as Tools


An internal support agent had 23 tools registered. Someone asked it: "How many orders did customer 8842 place last month?"

It called search_documents("customer 8842 orders"). That tool searches the help-centre knowledge base — articles about how to place an order, how to cancel one, what the refund window is. It returned three articles. The agent read them, found no number, and answered: "I could not find order history for that customer."

The right tool was run_query, which runs read-only SQL against the orders database and would have answered in 40 milliseconds. Why did the agent not pick it?

Python
search_documents:  "Search for information."run_query:         "Query the database."

Read those two descriptions as the model sees them — with no knowledge of your architecture, no idea that "documents" means help articles, no idea what is in "the database". The question contains the word "orders". One tool mentions searching for information, which is what the user wants. The model picked reasonably from what it was given.

The bug was not in the model. It was in two lines of text a developer wrote in four seconds.

Twenty-three tools, and the one it must pickWhich toolanswers this?Orders — lookup,history, statusPayments — charge, refundCustomers —profile, addressesShipping — track, rescheduleTickets — open,comment, close
Every registered tool costs tokens on every call and adds one more wrong option — grouping is what keeps selection accurate.

What a tool actually is

A tool is a function the model may ask you to call. That is the entire mechanism, and it is worth being precise about who does what:

Text
1. You send:      the conversation + a list of tool SCHEMAS2. Model returns: "call get_order(order_id='ORD-4471')"3. YOUR CODE:     actually calls get_order('ORD-4471')4. You send:      the conversation + the return value5. Model returns: an answer, or another tool call

The model never executes anything. It emits a structured request, and your code decides whether to honour it. Everything about tool design follows from that division: the model's only information about a tool is its schema, and your code's only protection against a bad request is validation.

The model does not see your function. It sees your description of it. If those two disagree, the description wins — and the failure appears in your code, not in the prompt.

Anatomy of a tool definition

Here is the same tool written badly and then properly.

Bad

JSON
{  "name": "get_data",  "description": "Gets data",  "parameters": {    "type": "object",    "properties": {      "query": { "type": "string" },      "opts":  { "type": "object" }    }  }}

Four separate failures. The name says nothing. The description says less. query could be SQL, a search phrase, a customer ID, or a question in English — the model will guess, and different calls will guess differently. And opts as a free-form object is an invitation to invent keys that your code has never heard of.

Good

JSON
{  "name": "get_customer_orders",  "description": "Return the order history for one customer from the orders database. Use this for questions about what a specific customer bought, when, and for how much. Does NOT search help articles or product descriptions - use search_help_articles for those. Returns at most 100 orders, newest first.",  "parameters": {    "type": "object",    "properties": {      "customer_id": {        "type": "integer",        "description": "Numeric internal customer ID, e.g. 8842. Not an email address and not an order ID."      },      "start_date": {        "type": "string",        "description": "Inclusive start of the range, ISO format YYYY-MM-DD, e.g. 2026-07-01."      },      "end_date": {        "type": "string",        "description": "Inclusive end of the range, ISO format YYYY-MM-DD."      },      "status": {        "type": "string",        "enum": ["placed", "shipped", "delivered", "cancelled", "any"],        "description": "Filter by order status. Use 'any' for all statuses."      }    },    "required": ["customer_id", "start_date", "end_date"]  }}

Every addition is load-bearing:

ElementPrevents
get_customer_orders not get_dataGuessing what the tool is for from the description alone
"Use this for questions about what a specific customer bought"Not recognising this tool applies to the question
"Does NOT search help articles — use search_help_articles"The exact confusion that produced the failure above
"Returns at most 100 orders"Concluding a customer has exactly 100 orders
customer_id is integerPassing "customer 8842" as a string
"Not an email address and not an order ID"Passing "ORD-4471" into a customer field
ISO format spelled out with an example"last month", "07/2026", "July"
enum on statusInventing "in transit" or "complete"
required listOmitting the date range and pulling all history

The negative statement — "does NOT do X, use Y for that" — is the single highest-value sentence you can add when two tools are adjacent in meaning. It is worth writing it on both tools, in both directions.

Constraints the schema cannot express

JSON Schema can say "integer" and "one of these five values". It cannot say "end_date must be after start_date", "the range must be under 366 days", or "this customer must belong to the caller's organisation". Those go in two places, and you need both.

In the description, so the model can get it right first time:

Text
"The range must be at most 366 days and end_date must not be before start_date. Requests outside these limits are rejected."

In the code, so a wrong call fails safely and informatively:

Python
from datetime import dateMAX_RANGE_DAYS = 366def get_customer_orders(customer_id, start_date, end_date, status="any"):    try:        start = date.fromisoformat(start_date)        end = date.fromisoformat(end_date)    except ValueError:        return {"error": "invalid_date_format",                "message": "Dates must be YYYY-MM-DD, e.g. 2026-07-01.",                "received": {"start_date": start_date,                             "end_date": end_date}}    if end < start:        return {"error": "invalid_range",                "message": (f"end_date {end_date} is before start_date "                            f"{start_date}. Swap them.")}    span = (end - start).days    if span > MAX_RANGE_DAYS:        return {"error": "range_too_large",                "message": (f"Requested {span} days; maximum is "                            f"{MAX_RANGE_DAYS}. Split into "                            f"{-(-span // MAX_RANGE_DAYS)} calls.")}    rows = db.fetch_orders(customer_id, start, end, status)    return {"customer_id": customer_id,            "order_count": len(rows),            "truncated": len(rows) == 100,            "orders": rows}

Two things about those error returns. They are returned, not raised, so they arrive in the agent's observation channel where the model can read them. And each one says what to do next — swap the dates, split into two calls. An error message that only reports what went wrong makes the agent guess; one that names the fix makes it succeed on the retry.

The truncated flag matters just as much. Without it, an agent that receives exactly 100 orders will confidently report "this customer placed 100 orders". With it, the agent knows the answer is a floor, not a count.

The function-calling round trip

Python
import json, osfrom anthropic import Anthropicclient = Anthropic()MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5")   # current default; set per deployment# Anthropic's API calls the schema field input_schema rather than parameters.TOOLS = [{"name": s["name"], "description": s["description"],          "input_schema": s["parameters"]}         for s in (GET_CUSTOMER_ORDERS_SCHEMA, SEARCH_HELP_ARTICLES_SCHEMA)]REGISTRY = {"get_customer_orders": get_customer_orders,            "search_help_articles": search_help_articles}messages = [{"role": "user",             "content": "How many orders did customer 8842 place "                        "in July 2026?"}]while True:    resp = client.messages.create(        model=MODEL,        max_tokens=1024,        tools=TOOLS,        messages=messages,    )    messages.append({"role": "assistant", "content": resp.content})    if resp.stop_reason != "tool_use":        print("".join(b.text for b in resp.content if b.type == "text"))        break    results = []    for block in resp.content:        if block.type != "tool_use":            continue        fn = REGISTRY.get(block.name)        if fn is None:            out = {"error": "unknown_tool",                   "message": f"No tool named {block.name}. "                              f"Available: {list(REGISTRY)}"}        else:            try:                out = fn(**block.input)            except TypeError as e:                out = {"error": "bad_arguments", "message": str(e)}            except Exception as e:                out = {"error": type(e).__name__, "message": str(e)}        results.append({"type": "tool_result",                        "tool_use_id": block.id,                        "content": json.dumps(out)})    messages.append({"role": "user", "content": results})

Four details that are easy to get wrong:

  • The assistant turn is appended before results. The tool-use block must remain in the conversation or the model loses track of what it asked for.
  • tool_use_id must match. When a model requests three tools in one turn, this is what pairs each result with its request.
  • All results go back in one user message. Not three separate turns.
  • Every failure path returns a dict. Unknown tool, wrong arguments, tool crash — all become observations, never exceptions that kill the loop.

The schemas in this lesson use the provider-neutral parameters key. Each provider wraps the same JSON Schema slightly differently — Anthropic's Messages API wants it as input_schema, OpenAI's function tools as parameters — which is why the code converts once, at the edge, instead of scattering provider formats through the tool library.

Organising a tool library

Past about eight tools, an unstructured list stops working. Three problems arrive together: the model picks wrong more often, the schemas eat your context window, and nobody on the team can remember what exists.

The token arithmetic

A well-written schema costs roughly 150–200 tokens. With 23 tools that is about 4,200 tokens re-sent on every model call. Over a 10-step agent run:

Tools in promptSchema tokens/callOver 10 stepsCost at an example 3 USD per million input
5~9009,0002.7 cents
12~2,20022,0006.6 cents
23~4,20042,00012.6 cents
60~11,000110,00033 cents

Thirty-three cents per run of pure overhead, before the agent has done anything, and multiplied across thousands of runs a day. Prompt caching reduces the price but not the context consumed.

The selection problem

The accuracy cost is worse than the token cost. Tool selection is a discrimination task: the model must pick one option from N descriptions. Every additional tool adds another chance for a near-miss, and near-misses cluster — a library with search_documents, search_articles, find_content and lookup_kb is four tools that are mutually confusable no matter how good the model is.

Three fixes, in order of how much they help:

Merge overlapping tools. Those four search tools are one tool with a source enum. One schema, one decision, no ambiguity:

JSON
{  "name": "search",  "description": "Full-text search over one content source.",  "parameters": {    "type": "object",    "properties": {      "source": {        "type": "string",        "enum": ["help_articles", "product_catalogue",                 "internal_wiki", "past_tickets"],        "description": "help_articles: customer-facing how-to guides. product_catalogue: SKUs, specs, prices. internal_wiki: engineering and policy docs. past_tickets: resolved support conversations."      },      "query": { "type": "string", "description": "Search phrase." }    },    "required": ["source", "query"]  }}

Namespace by domain. orders_get, orders_cancel, billing_refund, billing_invoice. The prefix carries information before the model reads a single word of description, and it makes the library legible to humans too.

Load tools dynamically. Keep a compact index of all tools, embed the descriptions, and at each step retrieve only the 5–8 most relevant to the current goal:

Python
class ToolLibrary:    def __init__(self, tools, embedder):        self.tools = {t["name"]: t for t in tools}        self.vectors = {t["name"]: embedder(t["description"])                        for t in tools}        self.embedder = embedder    def relevant(self, goal, k=6, always=("finish",)):        q = self.embedder(goal)        ranked = sorted(self.tools,                        key=lambda n: -cosine(q, self.vectors[n]))        chosen = list(always) + [n for n in ranked if n not in always][:k]        return [self.tools[n] for n in chosen]

The always parameter is not decoration. Tools the agent needs regardless of topic — finishing, escalating, asking the user — must never be retrieved away, or the agent will occasionally be unable to stop.

Composing tools

Most real work chains tools, where one output becomes another's input. This only works if the types line up.

Text
get_customer_orders(8842, ...)  ->  { orders: [ {order_id: "ORD-4471",                                                 total_paise: 425000,                                                 status: "delivered"} ] }                                            │                                            ▼  order_idget_shipment(order_id="ORD-4471")   ->  { carrier: "Delhivery",                                          tracking: "1Z994A",                                          delivered_at: "2026-07-19" }                                            │                                            ▼  trackingget_tracking_events(tracking="1Z994A")

The chain holds because get_customer_orders returns order_id in exactly the format get_shipment expects. Break that and the agent must transform values itself — which it will do by guessing, and guessing about identifier formats fails silently.

Composition problemSymptomFix
Tool A returns id: 4471, tool B wants "ORD-4471"404 on every chained callReturn the canonical form both tools use
A returns cents, B expects rupeesAmounts off by 100×Put the unit in the field name everywhere
A returns a list, B takes one itemAgent loops one-by-one, burning stepsGive B a batch parameter
A returns 40KB of JSONContext fills; the useful field is buriedAdd a fields parameter; return summaries by default
A's date is 2026-07-19, B wants epoch secondsSilently wrong time windowsISO 8601 strings at every boundary

Where a chain is fixed and frequent, collapse it into one tool. If "customer ID to delivery status" is asked a hundred times a day, write get_delivery_status(customer_id) that runs all three internally. You turn three model calls into one, remove two chances to pick the wrong tool, and cut the token cost roughly threefold. The agent's judgement should be spent on decisions that are genuinely open, not on re-deriving a fixed sequence.

Every step you can move from the agent's reasoning into a tool's implementation is a step that becomes fast, cheap, testable and deterministic.

Real tool groupings

AgentRead toolsWrite toolsBoundary to enforce
Supportorders_get, billing_invoices, search(source)billing_refund, tickets_escalate, email_sendRefunds capped and logged; email templates fixed, not free text
Researchweb_search, page_fetch, pdf_extractnotes_appendDomain allowlist; fetch size cap; no outbound POSTs
Data analysissql_select, schema_describe, chart_render— none —Read-only DB role; statement timeout; row cap
Codingfile_read, grep, tests_runfile_write, git_commitWrites confined to the repo; no push; no network in tests
Opsmetrics_query, logs_search, service_statusservice_restart, scale_setBlast-radius limit; approval for anything customer-facing

The pattern across every row: read tools far outnumber write tools, and every write tool has a named constraint. That asymmetry is deliberate. Reads are cheap to get wrong and writes are not, so give the agent generous read access and grudging write access.

Where people get this wrong

Writing the description for yourself. "Gets order data" is meaningful to you because you know what an order is here. Write for a competent stranger who has your schema and nothing else. If two people on your team would interpret the description differently, so will the model.

Exposing raw API responses. A payment provider returning 60 fields of which 3 matter costs you tokens on every call and hides the signal. Wrap it, return what the agent needs, and add a fields parameter for the rare case where it needs more.

Raising exceptions instead of returning errors. An exception ends the loop. A returned error dict keeps the agent alive and gives it something to act on. The only things that should raise are conditions where continuing would be unsafe.

One tool that does everything. The mirror image of too many tools: a single execute(command) that takes a natural-language instruction. It looks elegant and it is untestable, unvalidatable, and impossible to permission — you cannot grant read-only access to a tool whose scope is "anything".

Trusting arguments because a model produced them. Model output is untrusted input, and it can be influenced by text in the conversation. Validate types, ranges, ownership and permission on every call, exactly as you would for a request arriving from the internet — because in effect it is one.

Silent truncation. A tool that caps results at 100 and does not say so turns a limit into a lie. Any cap, filter or approximation must appear in the return value.

What this means when you build one

Write the tool schema before the implementation. Naming the tool, describing when to use it and when not to, and enumerating the parameter values forces you to decide what the tool is for — and that decision is the one the model inherits.

Then run the cheapest useful test there is: put your tool list in front of a colleague who does not know the system, give them ten real user questions, and ask which tool they would call. Every disagreement is a description bug, and you have found it for the price of ten minutes rather than a week of confused production traces.

Keep the active tool list under about eight. Merge near-duplicates behind an enum, namespace what remains by domain, and retrieve dynamically if the full library is large — pinning the tools that must always be available.

Make every return value structured, unit-suffixed, and honest about its own limits: truncated, as_of, excluded, approximate. And write your error returns as instructions. The difference between {"error": "invalid_range"} and {"error": "invalid_range", "message": "end_date 2026-07-01 is before start_date 2026-07-31. Swap them."} is the difference between an agent that fails and an agent that fixes itself on the next turn — which is, in the end, what you were building the loop for.