Course Content
Structured Output and Function Calling
3 sections · 6 lessons
Mini Project: Function-Calling Chatbot
You wire the loop together in twenty minutes. The first question works — "what's the weather in Tokyo?" comes back with a real temperature. Delighted, you type the follow-up, "and the forecast?", and the process dies:
BadRequestError: 400 - An assistant message with 'tool_calls' must befollowed by tool messages responding to each tool_call_id. The followingtool_call_ids did not have response messages: call_9QkR2mVb8xNfLpNothing is wrong with your API key, your schema, or the weather service. The bug is in the message history. On the first turn you appended the tool result but forgot to append the assistant message that requested it, so the history now contains an orphaned answer to a question the transcript never records being asked. The API refuses it, correctly.
That is the character of this whole build. The clever part — the model deciding which tool to call — takes about fifteen lines. The rest is bookkeeping: keeping a message array consistent, validating arguments you did not write, capping loops, caching, and failing in a way a user can understand. Getting that bookkeeping right is the actual skill.
What you are building
A command-line weather assistant that holds a multi-turn conversation, decides for itself when it needs live data, calls a real API to get it, and answers in plain English. Concretely:
| Tool | Arguments | Returns | Called when the user asks |
|---|---|---|---|
get_current_weather | location, units | 6 shaped fields | "Is it raining in Leeds?" |
get_forecast | location, days (1–5) | Per-day min/max/condition | "Will I need a coat this week?" |
Two tools is deliberate. One tool teaches you nothing about tool selection; ten tools buries the loop in plumbing. Two is the smallest number where the model has to make a real choice, and where you can watch it choose wrongly and fix it.
What the finished thing must do that a toy version does not: survive an unknown city, survive the weather service being down, refuse to loop forever, never leak an API key into a log, and answer a follow-up question without re-asking which city you meant.
Setting up
python -m venv .venv && source .venv/bin/activatepip install openai requests python-dotenv jsonschema pytestYou need two credentials: a model API key, and a free OpenWeatherMap key (the free tier allows 60 calls per minute, which is far more than this project needs once caching is in).
1# .env - and add .env to .gitignore before you write anything into it2OPENAI_API_KEY=sk-...3OPENAI_MODEL=gpt-6-sol # any current model that supports tool calling4WEATHER_API_KEY=...1import os2from dotenv import load_dotenv34load_dotenv()5WEATHER_API_KEY = os.environ["WEATHER_API_KEY"] # KeyError at start-up, not6OPENAI_API_KEY = os.environ["OPENAI_API_KEY"] # a confusing 401 an hour in7MODEL = os.environ["OPENAI_MODEL"] # change models without a deployUse os.environ["..."] rather than os.getenv("..."). The first fails immediately and loudly if the key is missing; the second returns None, which travels silently into an HTTP request and surfaces forty minutes later as an inscrutable 401 Invalid API key from a service you were not thinking about.
Declaring the tools
A tool declaration is three things: a name, a description written for the model, and a JSON Schema for the arguments. The description is not documentation — it is the entire basis on which the model decides whether this tool answers the user's question. Vague descriptions produce wrong tool choices, and no amount of system-prompt scolding fixes that.
1[2 {3 "type": "function",4 "function": {5 "name": "get_current_weather",6 "description": "Get conditions RIGHT NOW for one location. Use for questions about the present ('is it raining', 'how hot is it'). Do not use for future days.",7 "parameters": {8 "type": "object",9 "properties": {10 "location": {11 "type": "string",12 "description": "City, optionally with ISO country code: 'Leeds, GB'"13 },14 "units": {15 "type": "string",16 "enum": ["metric", "imperial"],17 "description": "metric = Celsius, imperial = Fahrenheit. Default metric."18 }19 },20 "required": ["location"],21 "additionalProperties": false22 }23 }24 },25 {26 "type": "function",27 "function": {28 "name": "get_forecast",29 "description": "Get a day-by-day forecast for the NEXT 1-5 days. Use for any question about tomorrow or later. Cannot return more than 5 days.",30 "parameters": {31 "type": "object",32 "properties": {33 "location": {"type": "string"},34 "days": {"type": "integer", "minimum": 1, "maximum": 5}35 },36 "required": ["location"],37 "additionalProperties": false38 }39 }40 }41]Four choices in there are doing real work, and each one exists because of a specific failure:
- "RIGHT NOW" versus "NEXT 1-5 days" in the descriptions. Without an explicit contrast the model calls
get_current_weatherfor "what's it like tomorrow" perhaps one time in five, and then confidently reports today's temperature as tomorrow's. enumonunits. A free-text string gets you"metric","celsius","C", and"Metric"across four calls, and your handler has to guess.maximum: 5ondays. The model does not know your upstream stops at five days unless the schema says so — and it will cheerfully ask for 10.additionalProperties: false. Without it, an invented argument like"detailed": truelands in your handler's keyword arguments and raisesTypeError: get_forecast() got an unexpected keyword argument 'detailed'in the middle of the loop.
A tool description is not documentation for humans who might read it later. It is the entire evidence the model has when deciding whether this tool answers the question in front of it.
Writing the handlers
Every handler obeys the same three rules: it validates what it was given, it never raises, and it returns the smallest useful result rather than the upstream's raw body.
1import time, requests23GEO_URL = "https://api.openweathermap.org/geo/1.0/direct"4BASE_URL = "https://api.openweathermap.org/data/2.5"5_cache: dict[tuple, tuple[float, dict]] = {} # key -> (expires_at, payload)67def _cached(key: tuple, ttl_s: int, produce):8 now = time.monotonic()9 hit = _cache.get(key)10 if hit and hit[0] > now:11 return hit[1]12 value = produce()13 if value.get("ok"): # never cache a failure14 _cache[key] = (now + ttl_s, value)15 return value1617def _geocode(location: str) -> dict:18 def fetch():19 r = requests.get(GEO_URL, timeout=(3, 5), params={20 "q": location, "limit": 1, "appid": WEATHER_API_KEY})21 if r.status_code != 200:22 return {"ok": False, "error": "geocode_failed"}23 rows = r.json()24 if not rows:25 return {"ok": False, "error": "location_not_found",26 "message": f"No place matched {location!r}. "27 f"Try 'City, CC', e.g. 'Cambridge, GB'."}28 return {"ok": True, "lat": rows[0]["lat"], "lon": rows[0]["lon"],29 "label": f"{rows[0]['name']}, {rows[0].get('country', '')}"}30 return _cached(("geo", location.lower()), 86_400, fetch) # places don't move3132def get_current_weather(location: str, units: str = "metric") -> dict:33 if not isinstance(location, str) or not location.strip():34 return {"ok": False, "error": "invalid_location"}35 place = _geocode(location)36 if not place["ok"]:37 return place3839 def fetch():40 try:41 r = requests.get(f"{BASE_URL}/weather", timeout=(3, 5), params={42 "lat": place["lat"], "lon": place["lon"],43 "units": units, "appid": WEATHER_API_KEY})44 except requests.RequestException:45 return {"ok": False, "error": "weather_service_unreachable"}46 if r.status_code != 200:47 return {"ok": False, "error": "weather_service_error",48 "message": f"HTTP {r.status_code}"}49 d = r.json()50 return {"ok": True, "location": place["label"],51 "temperature": round(d["main"]["temp"], 1),52 "feels_like": round(d["main"]["feels_like"], 1),53 "condition": d["weather"][0]["description"],54 "humidity_pct": d["main"]["humidity"],55 "units": "C" if units == "metric" else "F"}5657 return _cached(("now", place["label"], units), 600, fetch) # 10-minute TTLThe result the model actually receives is small on purpose:
1{2 "ok": true,3 "location": "Tokyo, JP",4 "temperature": 22.4,5 "feels_like": 21.1,6 "condition": "broken clouds",7 "humidity_pct": 65,8 "units": "C"9}Note "units": "C" travelling back with the numbers. Without it the model has a bare 22.4 and has to infer the scale from context, which it will occasionally get wrong and report as 22 degrees Fahrenheit. Any number that has a unit should carry its unit.
What the 10-minute cache is worth
Weather does not change meaningfully in ten minutes, so a short time-to-live is free accuracy-wise and dramatic cost-wise. Take a realistic burst: 300 questions arrive in ten minutes, spread across 12 cities. Without the cache that is 300 geocode calls plus 300 weather calls. With it, 12 weather calls (one per city per window) and roughly zero geocode calls after the first day, since the coordinates of Leeds are cached for 24 hours. That is a 96 per cent reduction in upstream traffic, and it takes the free tier's 60-calls-per-minute limit from "constantly breached" to "never approached".
The if value.get("ok") guard matters as much as the cache itself. Caching a failure means one transient upstream blip becomes ten minutes of confidently reported failure for every user asking about that city.
The conversation loop
The loop uses the Chat Completions message format from the function-calling lesson; on the Responses API the same invariant holds with function_call items and function_call_output entries paired by call_id. The loop is a small state machine over a growing list of messages. Its one hard rule: every assistant message containing tool_calls must be followed by exactly one tool message per call, each carrying the matching tool_call_id. Violate it and you get the 400 that opened this lesson.
1import json2import jsonschema3from openai import OpenAI45client = OpenAI(api_key=OPENAI_API_KEY)6HANDLERS = {"get_current_weather": get_current_weather, "get_forecast": get_forecast}78SYSTEM = ("You are a weather assistant. Use the tools for anything about actual "9 "conditions; never guess a temperature. If a tool returns ok:false, "10 "explain the problem plainly and suggest what the user could try.")1112class WeatherBot:13 def __init__(self, max_tool_rounds: int = 4):14 self.messages = [{"role": "system", "content": SYSTEM}]15 self.max_tool_rounds = max_tool_rounds1617 def ask(self, user_input: str) -> str:18 self.messages.append({"role": "user", "content": user_input})1920 for _ in range(self.max_tool_rounds):21 reply = client.chat.completions.create(22 model=MODEL, messages=self.messages,23 tools=TOOLS, tool_choice="auto",24 ).choices[0].message2526 # 1. ALWAYS append the assistant message, tool calls and all27 self.messages.append(reply.model_dump(exclude_none=True))2829 if not reply.tool_calls:30 return reply.content3132 # 2. One tool message per tool call, in the same order33 for call in reply.tool_calls:34 result = self._run(call.function.name, call.function.arguments)35 self.messages.append({36 "role": "tool",37 "tool_call_id": call.id, # the bit people forget38 "content": json.dumps(result),39 })4041 return ("I got stuck looking that up after several attempts. "42 "Could you rephrase the question?")4344 def _run(self, name: str, raw_args: str) -> dict:45 handler = HANDLERS.get(name)46 if handler is None:47 return {"ok": False, "error": "unknown_tool", "message": f"No tool {name}."}48 try:49 args = json.loads(raw_args)50 except json.JSONDecodeError:51 return {"ok": False, "error": "unparseable_arguments",52 "message": "Arguments were not valid JSON. Try again."}53 try:54 jsonschema.validate(args, SCHEMAS[name])55 except jsonschema.ValidationError as e:56 return {"ok": False, "error": "invalid_arguments", "message": e.message}57 try:58 return handler(**args)59 except Exception as e: # last line of defence60 log.exception("handler %s failed", name)61 return {"ok": False, "error": "tool_failed", "message": type(e).__name__}Why the loop is capped
max_tool_rounds=4 is not decoration. Models can get into a rut — call a tool, read a confusing result, call the same tool again with a trivially different argument, forever. Every round is a full API call whose input is the entire conversation so far, so the cost climbs as the history grows.
Suppose the first request is 700 tokens of input and each round adds roughly 250 tokens of tool call plus tool result. A 12-round loop sends 700, 950, 1,200, … 3,450 tokens — an arithmetic series summing to 12 × (700 + 3450) / 2 = 24,900 input tokens. At 3 dollars per million that is 7.5 cents, against 0.45 cents for a normal single-tool turn: roughly 17 times the cost, for a question that never gets an answer. The cap converts an unbounded failure into a polite one.
What the message array looks like after one question
This is the single most useful thing to print while debugging. After "What's the weather in Tokyo?", self.messages holds:
1[2 {"role": "system", "content": "You are a weather assistant..."},3 {"role": "user", "content": "What's the weather in Tokyo?"},4 {"role": "assistant", "content": null,5 "tool_calls": [6 {"id": "call_9QkR2mVb8xNfLp", "type": "function",7 "function": {"name": "get_current_weather",8 "arguments": "{\"location\":\"Tokyo\",\"units\":\"metric\"}"}}9 ]},10 {"role": "tool", "tool_call_id": "call_9QkR2mVb8xNfLp",11 "content": "{\"ok\":true,\"location\":\"Tokyo, JP\",\"temperature\":22.4,\"condition\":\"broken clouds\",\"humidity_pct\":65,\"units\":\"C\"}"},12 {"role": "assistant",13 "content": "It's 22.4C in Tokyo with broken clouds and 65% humidity."}14]Two details are worth staring at. The arguments field is a JSON string, not an object — hence json.loads before you can touch it, and hence the possibility, however rare with modern constrained decoding, of it not parsing at all. And the content of a tool message is also a string; whatever your handler returned must be serialised. Return a datetime or a Mongo ObjectId from a handler and json.dumps raises TypeError: Object of type datetime is not JSON serializable, killing the turn.
Parallel tool calls
Ask "compare London and Cairo" and one assistant message may contain two tool calls at once. The loop above already handles this — it iterates reply.tool_calls — but only because it appends one tool message per call. Code written for a single call, using reply.tool_calls[0], produces the orphaned-call 400 the moment a user asks a comparison question.
Append the assistant message before you run anything, and one tool message per tool call afterwards. Almost every mysterious 400 from a function-calling loop is a violation of that pairing.
When the model gets the arguments wrong
The schema says days is between 1 and 5. A user asks "what's the weather doing over the next week?" and the model, reasonably enough, emits:
{"name": "get_forecast", "arguments": "{\"location\":\"Bristol\",\"days\":7}"}Validation catches it before your handler runs, and the interesting part is what you do next. The wrong move is to raise, or to return "error". The right move is to hand the model a message it can act on:
{"ok": false, "error": "invalid_arguments", "message": "7 is greater than the maximum of 5"}Feed that back as the tool result and the loop's next round produces:
{"name": "get_forecast", "arguments": "{\"location\":\"Bristol\",\"days\":5}"}…followed by an answer that says "here are the next five days — that's as far ahead as I can see." The user gets a good answer, nobody sees a stack trace, and the self-correction cost one extra round trip. This is the payoff for making handlers total functions that return errors instead of raising them: the model is a retry mechanism, if you talk to it.
The same pattern covers a genuinely malformed response. Occasionally — far more often with older models or when the response is truncated by a low max_tokens — you get arguments like:
{"location": "Bristol", "days":json.loads raises JSONDecodeError: Expecting value: line 1 column 28 (char 27). Your _run catches it, returns unparseable_arguments, and the model reissues the call. Without that try/except the whole process falls over on a transient formatting glitch.
| What arrives | Caught by | Returned to the model | Typical outcome |
|---|---|---|---|
| Truncated JSON | json.loads | unparseable_arguments | Model reissues the call |
days: 7 | jsonschema | invalid_arguments + reason | Model retries with 5 |
"detailed": true | additionalProperties: false | invalid_arguments | Model drops the field |
| City that does not exist | Geocode returning no rows | location_not_found + hint | Model asks the user to clarify |
| Upstream 503 | status_code != 200 | weather_service_error | Model explains, does not invent |
Testing without touching the network
Tests that call the real weather API are slow, flaky, and quietly consume your quota; tests that call the real model are all of that plus non-deterministic. Split the test suite by what it actually exercises.
| Level | What is faked | What it proves |
|---|---|---|
| Handler unit tests | HTTP layer | Shaping, error mapping, caching, validation |
| Loop tests | The model client | Message pairing, iteration cap, tool dispatch |
| Schema tests | Nothing | Bad arguments are rejected, good ones accepted |
| Live smoke test | Nothing (run manually) | Keys work, upstream contract unchanged |
1def test_unknown_city_returns_structured_error(monkeypatch):2 monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(200, []))3 r = get_current_weather("Nowhereville")4 assert r["ok"] is False and r["error"] == "location_not_found"5 assert "Try 'City, CC'" in r["message"] # the hint the model needs67def test_failures_are_not_cached(monkeypatch):8 monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(503, {}))9 get_current_weather("Leeds")10 monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(200, TOKYO_JSON))11 assert get_current_weather("Leeds")["ok"] is True # recovers immediately1213def test_days_above_maximum_is_rejected():14 bot = WeatherBot()15 r = bot._run("get_forecast", '{"location":"Bristol","days":7}')16 assert r["error"] == "invalid_arguments" and "maximum of 5" in r["message"]1718def test_loop_stops_at_the_cap():19 bot = WeatherBot(max_tool_rounds=3)20 with always_calls_a_tool(): # fake client, never finishes21 answer = bot.ask("weather?")22 assert "stuck" in answer23 assert sum(1 for m in bot.messages if m["role"] == "tool") == 32425def test_every_tool_call_has_a_matching_tool_message():26 bot = WeatherBot()27 with scripted_client([two_parallel_calls(), plain_answer()]):28 bot.ask("compare London and Cairo")29 ids = [c["id"] for m in bot.messages if m.get("tool_calls")30 for c in m["tool_calls"]]31 responses = [m["tool_call_id"] for m in bot.messages if m["role"] == "tool"]32 assert sorted(ids) == sorted(responses)That last test is the one that earns its keep. It encodes the invariant behind the 400 at the top of this lesson, and it catches a regression the moment someone "simplifies" the loop to handle a single tool call.
Cost and latency, in numbers
A single tool-using turn makes two model calls, not one, plus the upstream fetch. Timed end to end on a cold cache:
| Stage | Time | Tokens |
|---|---|---|
| Model call 1 (decide + emit tool call) | 900 ms | ~700 in, ~40 out |
| Geocode (cache miss) | 180 ms | — |
| Weather fetch | 350 ms | — |
| Model call 2 (read result, write answer) | 1,400 ms | ~800 in, ~90 out |
| Total | 2,830 ms | 1,500 in, 130 out |
At 3 dollars per million input tokens and 15 per million output, that is 1500/1e6 × 3 = 0.0045 plus 130/1e6 × 15 = 0.00195, so roughly 0.65 cents per question — about 6.50 dollars per thousand questions. On a warm cache the geocode disappears and the total drops to 2,650 ms.
The useful observation is where the time goes: 2,300 of those 2,830 milliseconds are the two model calls. Optimising your weather fetch from 350 ms to 200 ms is invisible to the user. Streaming the second model call so text appears as it is generated is not — it takes perceived latency from "nearly three seconds of nothing" to "about one second, then words".
When it breaks in front of a user
| Symptom | Cause | Fix |
|---|---|---|
tool_call_ids did not have response messages | Assistant message not appended, or only the first tool call answered | Append the assistant message first; loop over all tool_calls |
TypeError: got an unexpected keyword argument | Invented argument reached handler(**args) | additionalProperties: false plus schema validation before dispatch |
| Reports today's weather for "tomorrow" | Tool descriptions do not contrast present with future | Say "RIGHT NOW" and "NEXT 1-5 days" explicitly |
| Invents a temperature, no tool call in the log | tool_choice left unset, or the system prompt permits guessing | tool_choice="auto" plus "never guess a temperature" |
| Same city, different answers minutes apart | No cache; upstream rounding differs per call | Cache successful lookups for 10 minutes |
| Slow, then a 401 an hour into a session | os.getenv returned None for a missing key | os.environ["..."] so it fails at start-up |
| Conversation gets slower every turn | History grows without bound; every turn resends it all | Trim or summarise older turns above a token budget |
Log the tool name, the arguments and the outcome for every call, tagged with a conversation id. Without that log, "it gave a weird answer yesterday" is unanswerable; with it, it is a two-minute lookup.
Where to take it
The interesting extensions are the ones that change the shape of the system rather than adding another endpoint:
- Persistence. Store
messagesper user in SQLite or Redis so a session survives a restart. This immediately forces the history-trimming question, because a stored conversation grows forever and every turn pays for the whole thing. - A second, unrelated data source — air quality, or tide times. Two tools from one API is an easy choice for the model; four tools across two domains is where descriptions start to matter, and where you will see your first genuinely wrong tool selection.
- An HTTP front end. Wrap
ask()in a FastAPI endpoint keyed by session id. Now concurrency is real: the module-level_cachedictionary is shared across requests, which is fine for reads and wrong the moment you add anything stateful. - Streaming. The single biggest perceived-speed win available, worth more than any backend optimisation.
- A tool that writes something — "text me if it rains tomorrow". The moment a tool has a side effect, everything changes: it needs idempotency so a retried loop does not send two texts, and it needs a per-conversation cap so a confused model cannot send forty.
Build the read-only version first and get the loop invariants right, because they do not change when the tools get more interesting. The bookkeeping you have just written — append the assistant message, answer every tool call by id, validate arguments, return errors instead of raising, cap the rounds, cache the reads, log everything — is the same bookkeeping behind an agent that manages calendars or files expenses. Only the handlers change.