Course Content
Structured Output and Function Calling
3 sections · 6 lessons
Function Calling and Tool Use
A retailer puts a support assistant on its order-status page. A customer types "where is order A-99213?" and gets back:
Your order A-99213 shipped on Tuesday 3rd March via Royal Mailand should arrive within 2-3 working days. Tracking numberRM884120395GB.Every detail is invented. There is no order A-99213 in the system, no Royal Mail booking, and that tracking number belongs to nobody. The model was asked a question about the world, had no access to the world, and did the only thing a text generator can do: produced the most plausible-sounding continuation. Plausible order updates look exactly like that.
No prompt fixes this. "Only answer from real data" does not help when there is no real data in the context. The model needs a way to reach the order system — and, crucially, a way to do so that your code can act on rather than read. That mechanism is function calling, also called tool use.
The trick underneath is small and worth stating up front: the model never runs anything. It emits a structured, schema-shaped request — "call get_order_status with {"order_id": "A-99213"}" — and your code decides whether to honour it. Function calling is schema-constrained generation pointed at a different question. Instead of "what does this data look like", the question is "which capability should be invoked, and with what arguments".
The loop, end to end
1. You declare the tools available, each as a name, a description, and a JSON Schema for its arguments: get_order_status(order_id) get_weather(location, date)2. The user asks something: "Where is order A-99213?"3. The model reads the tools plus the message. Instead of guessing an answer, it emits a tool call and stops: name: get_order_status input: {"order_id": "A-99213"} The response carries a marker saying "I am waiting on a tool", not "I have finished".4. YOUR code validates the call, looks up the real function, runs it: get_order_status("A-99213") -> {"status": "in_transit", "carrier": "DPD", "eta": "2026-03-06", "tracking": "DPD9917342"}5. You append the model's tool call AND your result to the conversation, and call the model again.6. The model reads the real data and writes the final answer: "Order A-99213 is in transit with DPD and is due on 6 March. Tracking: DPD9917342."Two things about step 3 are load-bearing. First, the model stops mid-turn — the response has a distinct finish signal (tool_calls on OpenAI, stop_reason: "tool_use" on Anthropic) and your loop must branch on it rather than assuming every response is an answer. Second, the arguments came out of a schema-constrained decoder, so they are the right shape by construction. What they are not is trustworthy: a well-formed {"order_id": "A-00001"} for somebody else's order is exactly as well-formed as the right one.
The model proposes; your code disposes. Every safety property of a tool-using system lives in the gap between the proposal and the execution, and that gap is entirely yours to defend.
A note on API generations
Plenty of tutorials still show a functions=[...] parameter paired with function_call="auto", and responses carrying a single message.function_call. That was OpenAI's first version of this feature and it was deprecated in 2023. It was replaced by tools=[...] with tool_choice in the Chat Completions API, and OpenAI's newer Responses API, which it recommends for new projects, flattens the shapes again.
| Deprecated (2023) | Chat Completions (still supported) | Responses API |
|---|---|---|
functions=[{name, description, parameters}] | tools=[{"type": "function", "function": {...}}] | tools=[{"type": "function", "name": ..., "parameters": ...}] |
function_call="auto" | tool_choice="auto" | tool_choice="auto" |
message.function_call (one only) | message.tool_calls (a list) | function_call items in response.output (a list) |
{"role": "function", "name": ...} | {"role": "tool", "tool_call_id": ...} | {"type": "function_call_output", "call_id": ...} |
The code in this lesson and in the mini project uses the Chat Completions shape, because its explicit message list makes the bookkeeping easy to see. The loop is the same in the Responses API: append the model's function_call items to your input list, then one function_call_output per call_id. One difference to know: in Chat Completions a function is non-strict unless you set strict, while the Responses API tries strict mode when you leave strict out.
The last row is the substantive one. Matching results to calls by name is ambiguous the moment the model requests the same tool twice in one turn — two get_weather calls for two cities have the same name and different arguments. IDs make the pairing unambiguous.
Step 1: declaring tools
A tool definition is a name, a description, and a JSON Schema for the arguments.
1TOOLS = [2 {3 "type": "function",4 "function": {5 "name": "get_order_status",6 "strict": True, # enforce the schema while decoding7 "description": (8 "Look up the current shipping status of a customer order. "9 "Use this whenever the user asks where an order is, when it "10 "will arrive, or whether it has shipped. Requires the order "11 "reference in the form A-99213."12 ),13 "parameters": {14 "type": "object",15 "properties": {16 "order_id": {17 "type": "string",18 "pattern": "^A-[0-9]{5}$",19 "description": "Order reference, e.g. 'A-99213'",20 }21 },22 "required": ["order_id"],23 "additionalProperties": False,24 },25 },26 },27 {28 "type": "function",29 "function": {30 "name": "start_return",31 "strict": True,32 "description": (33 "Open a return for a delivered order. Only call this after "34 "the user has explicitly asked to return something and has "35 "confirmed which order. Never call it speculatively."36 ),37 "parameters": {38 "type": "object",39 "properties": {40 "order_id": {"type": "string", "pattern": "^A-[0-9]{5}$"},41 "reason": {42 "type": "string",43 "enum": ["damaged", "wrong_item", "not_as_described",44 "changed_mind"],45 },46 },47 "required": ["order_id", "reason"],48 "additionalProperties": False,49 },50 },51 },52]The description is the whole ballgame
The model's only information about what a tool does is the text you wrote. It cannot read your implementation. A description of "gets order status" leaves it guessing about when the tool applies, what an order ID looks like, and whether it is safe to call twice. Compare:
| Weak description | What goes wrong | Strong description |
|---|---|---|
| "Gets order info" | Called for refund questions, billing questions, anything vaguely order-shaped | States exactly which user intents it serves |
| "Search the database" | Called constantly, for everything, because it sounds universal | Names the table and the kind of question it answers |
| "Send an email" | Fired without confirmation; the user gets mail they never asked for | "Only after the user has explicitly confirmed the recipient and content" |
| No format hint on an ID | Model invents #99213 or order-99213 | Gives the exact pattern and a concrete example |
Two more rules that save real debugging time. Keep the tool count small. Selection accuracy falls as the list grows, and the definitions cost input tokens on every single request. Eight tools at roughly 120 tokens each is 960 tokens per call; at 50,000 calls a day and 2.50 dollars per million input tokens that is 48 million tokens and 120 dollars a day, or about 3,600 dollars a month, just to remind the model what it can do. Put the tool block at the front of a cached prefix and most of that cost disappears. Avoid near-duplicate tools. get_order and fetch_order_details sitting side by side guarantees the model picks the wrong one some of the time, and no prompt fixes an ambiguity you built into the tool list.
Step 2: implementing the real functions
The functions are ordinary code. What makes them tool-shaped is that they take exactly the declared arguments, never raise into the loop, and always return something serialisable.
1import json, re23def get_order_status(order_id: str) -> dict:4 if not re.fullmatch(r"A-[0-9]{5}", order_id):5 return {"error": "invalid_order_id",6 "message": "Order references look like A-99213."}7 row = db.fetch_order(order_id)8 if row is None:9 return {"error": "not_found",10 "message": f"No order {order_id} exists."}11 return {12 "order_id": order_id,13 "status": row.status,14 "carrier": row.carrier,15 "eta": row.eta.isoformat(),16 "tracking": row.tracking,17 }1819TOOL_REGISTRY = {20 "get_order_status": get_order_status,21 "start_return": start_return,22}The registry is not decoration. It is the allow-list: a name the model invents that is not a key in that dict never reaches any code. Dispatching with globals()[name] or eval turns a hallucinated tool name into arbitrary code execution, and that is not a hypothetical.
Return errors, do not raise them
Note that get_order_status returns an error object instead of raising. This matters because an error is information the model can use. Hand it {"error": "not_found"} and it will tell the user the reference does not exist, or ask them to check it. Let the exception propagate and the whole turn dies with a stack trace the user never sees an explanation for.
| Failure | Return to the model | What the model does with it |
|---|---|---|
| Bad argument | {"error": "invalid_order_id", "message": "..."} | Asks the user for the correct reference |
| Not found | {"error": "not_found"} | Says so, offers to search another way |
| Upstream timeout | {"error": "timeout", "retryable": true} | Explains the delay; may retry once |
| Permission denied | {"error": "forbidden"} | Explains it cannot do that — without naming internals |
| Internal crash | {"error": "internal"} + log the trace server-side | Apologises; the stack trace stays out of the conversation |
That last row is a security boundary, not a style preference. Everything you return becomes conversation text the user may see. A raw traceback leaks file paths, library versions and sometimes connection strings.
Step 3: the request / execute / respond loop
1import json, os2from openai import OpenAI34client = OpenAI()5MODEL = os.environ.get("OPENAI_MODEL", "gpt-6-sol") # model ids live in config6MAX_ROUNDS = 578def chat(user_message: str) -> str:9 messages = [10 {"role": "system", "content": "You are an order support assistant."},11 {"role": "user", "content": user_message},12 ]1314 for _ in range(MAX_ROUNDS):15 response = client.chat.completions.create(16 model=MODEL,17 messages=messages,18 tools=TOOLS,19 tool_choice="auto",20 )21 msg = response.choices[0].message2223 if not msg.tool_calls:24 return msg.content # the model is done2526 messages.append(msg) # record the request verbatim2728 for call in msg.tool_calls:29 fn = TOOL_REGISTRY.get(call.function.name)30 if fn is None:31 result = {"error": "unknown_tool",32 "message": f"No tool named {call.function.name}."}33 else:34 try:35 args = json.loads(call.function.arguments)36 except json.JSONDecodeError:37 result = {"error": "bad_arguments"}38 else:39 try:40 result = fn(**args)41 except Exception:42 log.exception("tool %s failed", call.function.name)43 result = {"error": "internal"}4445 messages.append({46 "role": "tool",47 "tool_call_id": call.id,48 "content": json.dumps(result),49 })5051 return "I wasn't able to complete that request."The four bugs everyone writes at least once
Forgetting that arguments arrive as a string. This is the single most common one. The field looks like an object in the docs but it is a JSON-encoded string:
1{2 "id": "call_9fA2xQ",3 "type": "function",4 "function": {5 "name": "get_order_status",6 "arguments": "{\"order_id\": \"A-99213\"}"7 }8}Notice the escaped quotes. arguments is a string containing JSON, not a nested object. Calling fn(**call.function.arguments) without json.loads raises TypeError: argument after ** must be a mapping, not str. The reason for the encoding is that it is streamed in fragments and only becomes valid JSON once complete.
Not appending the assistant's tool-call message. If you run the function and append only the result, the conversation contains an answer to a question that was never asked. Providers reject this outright, with an error along the lines of "a message with role 'tool' must be a response to a preceding message with tool_calls".
Losing the ID pairing. Every tool_call_id must be echoed back, and every tool call must get a result — including the ones that failed. Skip a failed call and the request is malformed; results still owe an answer even when the answer is an error.
No iteration cap. Without MAX_ROUNDS, a model that keeps requesting the same failing tool will loop until you run out of money. Cap it, and log every run that hits the cap — a cap being hit regularly is a symptom of a tool description that does not match what the tool actually does.
Multiple and parallel tool calls
One assistant turn can request several tools at once. Ask "what's the weather in San Francisco and is Tesla up today?" and a single response may carry two tool_calls. That is a feature — the two lookups are independent, so run them concurrently:
1from concurrent.futures import ThreadPoolExecutor23with ThreadPoolExecutor(max_workers=4) as pool:4 futures = {5 call.id: pool.submit(dispatch, call)6 for call in msg.tool_calls7 }8results = {cid: f.result() for cid, f in futures.items()}910for call in msg.tool_calls: # preserve the original order11 messages.append({"role": "tool",12 "tool_call_id": call.id,13 "content": json.dumps(results[call.id])})Two 400 ms lookups run in parallel take 400 ms rather than 800 ms, and across a chain of six calls that is the difference between a two-second answer and a five-second one.
Two rules go with this. Return every result together — on Anthropic's API, all tool_result blocks must sit in a single user message; splitting them across several messages is accepted but quietly trains the model to stop making parallel calls. And parallel means independent: if tool B needs tool A's output, the model should be making them in separate turns, and if it is not, your descriptions have not made the dependency clear. Where you cannot tolerate concurrency at all, most providers offer a switch (parallel_tool_calls=False, or disable_parallel_tool_use) to force one call per turn.
Controlling the choice with tool_choice
| Setting | Behaviour | Use it when |
|---|---|---|
"auto" | Model decides: a tool, or a plain answer | Default. Conversational assistants |
"required" | Must call some tool; cannot answer in prose | A step where an answer without data is always wrong |
{"type": "function", "function": {"name": "x"}} | Must call that exact tool | Using the tool schema purely as an extraction format |
"none" | Cannot call anything; prose only | The summarising turn after the data is gathered |
The forced-single-tool case is worth dwelling on, because it is a technique rather than a setting. If you define one tool called record_ticket whose parameters are your ticket schema, and force it, you have built a structured extractor: the model has no path to a prose reply, so the only thing it can produce is a validated object. Before dedicated structured-output features existed, this was the way to get guaranteed shapes. Today, prefer the native structured-output mode from the previous lesson for plain extraction. Some of the newest models reject a forced tool choice with a 400 error, and on those the recommended replacement is exactly that: structured outputs, or auto with strict: true on the tool and a prompt that names it.
On Anthropic's API the same settings are written {"type": "auto"}, {"type": "any"} (must call some tool), {"type": "tool", "name": "x"} and {"type": "none"}.
The failure mode of "required" is worth naming too. Force a tool call on a turn where no tool is appropriate — the user said "thanks, that's all" — and the model will call something anyway, usually with invented arguments, because you removed its ability to decline. Use "required" on a specific step, not on a whole conversation.
The same idea across providers
Tool use is not one vendor's feature. The concepts map cleanly; the field names do not.
| Concept | OpenAI | Anthropic |
|---|---|---|
| Argument schema key | function.parameters | input_schema |
| "I want a tool" signal | finish_reason: "tool_calls" | stop_reason: "tool_use" |
| The request itself | message.tool_calls[] | A tool_use content block |
| Arguments format | JSON-encoded string | Already-parsed object in input |
| Returning a result | {"role": "tool", "tool_call_id": ...} | tool_result block inside a user message |
| Signalling a failed tool | Return an error object as content | is_error: true on the tool_result |
| Hard schema enforcement | strict: true in the function object | strict: true on the tool definition |
Here is one round trip on Anthropic's Messages API, as it sits in the messages list. The response that asked for the tool ended with stop_reason: "tool_use"; you append its content unchanged, then answer in a user message:
1[2 {"role": "assistant", "content": [3 {"type": "text", "text": "Let me check that order."},4 {"type": "tool_use", "id": "toolu_01A", "name": "get_order_status",5 "input": {"order_id": "A-99213"}}6 ]},7 {"role": "user", "content": [8 {"type": "tool_result", "tool_use_id": "toolu_01A",9 "content": "{\"status\": \"in_transit\", \"carrier\": \"DPD\"}"}10 ]}11]The pairing rule is the same as OpenAI's, with different names: every tool_use block needs a tool_result carrying its id as tool_use_id, and a failed tool still gets a result, marked "is_error": true.
Locally hosted open-weight models get there differently: many are fine-tuned to emit tool calls in a specific text format, which a runtime parses out and re-serialises into whichever API shape you asked for. The reliability varies far more than with hosted models, which makes client-side validation of the arguments non-negotiable rather than merely wise.
Write your dispatcher against your own internal call format and convert at the edges. The day you switch provider you will change one adapter instead of every handler.
Validating what the model asked for
Schema-constrained generation gives you well-formed arguments. It gives you nothing about whether those arguments are permissible. Consider a call your decoder is perfectly happy with:
1{2 "name": "start_return",3 "arguments": "{\"order_id\": \"A-00001\", \"reason\": \"damaged\"}"4}Correct types, correct pattern, valid enum value. It is also an order belonging to a different customer, and if your handler takes the ID at face value, you have just built an authorisation bypass driven by whatever text a user can get into the conversation.
So the dispatcher runs three checks, in this order, before anything executes:
- Is the tool real? The registry lookup. A name outside the dict is rejected, logged, and returned to the model as
unknown_tool. - Do the arguments validate? Re-check against the schema on your side. The provider may have ignored
pattern; you should not. - Is this caller allowed to do this, to this object? Not "is
A-00001a valid order" but "isA-00001this session's user's order". The session identity comes from your auth layer, never from the model's arguments.
1def dispatch(call, session):2 fn = TOOL_REGISTRY.get(call.function.name)3 if fn is None:4 return {"error": "unknown_tool"}56 try:7 args = json.loads(call.function.arguments)8 except json.JSONDecodeError:9 return {"error": "bad_arguments"}1011 ok, problems = validate_args(call.function.name, args)12 if not ok:13 return {"error": "invalid_arguments", "details": problems}1415 if "order_id" in args and not owns_order(session.user_id, args["order_id"]):16 log.warning("blocked cross-account access u=%s o=%s",17 session.user_id, args["order_id"])18 return {"error": "not_found"} # do not confirm it exists1920 return fn(**args)The not_found on an authorisation failure is deliberate. Replying "forbidden" tells the caller that order A-00001 exists, which is itself a leak; every probe becomes a confirmed hit.
What this changes about how you build
Adopting tool use quietly reorganises an application. Three consequences are worth planning for before you meet them.
Your tool surface is now a public API. The model is a client you cannot fully control, driven by text a user can influence. Every tool you expose is reachable by anyone who can type into the chat box, in any combination and any order. The question to ask of each new tool is not "is this useful" but "what is the worst sequence of calls someone could talk the model into". A read-only lookup and a delete endpoint deserve very different amounts of thought, and putting them behind the same dispatcher without different guards is how the interesting incidents happen.
Latency and cost are now multiplied by the number of rounds. One user question that needs three tool calls costs four model requests, plus three round trips to your own services. If each model call is 900 ms and each tool 300 ms, that is 4.5 seconds before the user sees a word — which is why streaming the final turn matters, why the tool block belongs in a cached prefix, and why fewer, more capable tools usually beat many granular ones.
Debugging moves from "read the code path" to "read the transcript". When a bot does something odd, the answer is almost always visible in the message list: which tools were offered, what the model asked for, what came back. Log the whole sequence — tool names, arguments, results, timings, and the round number — for every conversation, with a request ID that ties it to your normal service logs. The teams who skip this spend their incidents guessing, and a tool-use bug that you cannot reproduce is a bug you cannot fix.
A tool-using application is a conversation and a control plane wearing the same coat. Instrument the conversation as carefully as you instrument the control plane.