Structured Output and Function Calling

Course Content

Mini Project: Function-Calling Chatbot


You wire the loop together in twenty minutes. The first question works — "what's the weather in Tokyo?" comes back with a real temperature. Delighted, you type the follow-up, "and the forecast?", and the process dies:

Text
BadRequestError: 400 - An assistant message with 'tool_calls' must befollowed by tool messages responding to each tool_call_id. The followingtool_call_ids did not have response messages: call_9QkR2mVb8xNfLp

Nothing is wrong with your API key, your schema, or the weather service. The bug is in the message history. On the first turn you appended the tool result but forgot to append the assistant message that requested it, so the history now contains an orphaned answer to a question the transcript never records being asked. The API refuses it, correctly.

That is the character of this whole build. The clever part — the model deciding which tool to call — takes about fifteen lines. The rest is bookkeeping: keeping a message array consistent, validating arguments you did not write, capping loops, caching, and failing in a way a user can understand. Getting that bookkeeping right is the actual skill.

The message array after one weather questionsystem promptuser:weather Tokyoassistant:tool calltool: 18C, clearassistant: answerresendall of thisnext turnappendsDrop the tool-call and tool-result pair and the follow-up dies: the API rejects an orphaned result.
The array is the whole memory — every turn resends it, so appending in the wrong order is what breaks the loop.

What you are building

A command-line weather assistant that holds a multi-turn conversation, decides for itself when it needs live data, calls a real API to get it, and answers in plain English. Concretely:

ToolArgumentsReturnsCalled when the user asks
get_current_weatherlocation, units6 shaped fields"Is it raining in Leeds?"
get_forecastlocation, days (1–5)Per-day min/max/condition"Will I need a coat this week?"

Two tools is deliberate. One tool teaches you nothing about tool selection; ten tools buries the loop in plumbing. Two is the smallest number where the model has to make a real choice, and where you can watch it choose wrongly and fix it.

What the finished thing must do that a toy version does not: survive an unknown city, survive the weather service being down, refuse to loop forever, never leak an API key into a log, and answer a follow-up question without re-asking which city you meant.

Setting up

Bash
python -m venv .venv && source .venv/bin/activatepip install openai requests python-dotenv jsonschema pytest

You need two credentials: a model API key, and a free OpenWeatherMap key (the free tier allows 60 calls per minute, which is far more than this project needs once caching is in).

Bash
# .env  - and add .env to .gitignore before you write anything into itOPENAI_API_KEY=sk-...OPENAI_MODEL=gpt-6-sol          # any current model that supports tool callingWEATHER_API_KEY=...
Python
import osfrom dotenv import load_dotenvload_dotenv()WEATHER_API_KEY = os.environ["WEATHER_API_KEY"]   # KeyError at start-up, notOPENAI_API_KEY  = os.environ["OPENAI_API_KEY"]    # a confusing 401 an hour inMODEL           = os.environ["OPENAI_MODEL"]      # change models without a deploy

Use os.environ["..."] rather than os.getenv("..."). The first fails immediately and loudly if the key is missing; the second returns None, which travels silently into an HTTP request and surfaces forty minutes later as an inscrutable 401 Invalid API key from a service you were not thinking about.

Declaring the tools

A tool declaration is three things: a name, a description written for the model, and a JSON Schema for the arguments. The description is not documentation — it is the entire basis on which the model decides whether this tool answers the user's question. Vague descriptions produce wrong tool choices, and no amount of system-prompt scolding fixes that.

JSON
[  {    "type": "function",    "function": {      "name": "get_current_weather",      "description": "Get conditions RIGHT NOW for one location. Use for questions about the present ('is it raining', 'how hot is it'). Do not use for future days.",      "parameters": {        "type": "object",        "properties": {          "location": {            "type": "string",            "description": "City, optionally with ISO country code: 'Leeds, GB'"          },          "units": {            "type": "string",            "enum": ["metric", "imperial"],            "description": "metric = Celsius, imperial = Fahrenheit. Default metric."          }        },        "required": ["location"],        "additionalProperties": false      }    }  },  {    "type": "function",    "function": {      "name": "get_forecast",      "description": "Get a day-by-day forecast for the NEXT 1-5 days. Use for any question about tomorrow or later. Cannot return more than 5 days.",      "parameters": {        "type": "object",        "properties": {          "location": {"type": "string"},          "days": {"type": "integer", "minimum": 1, "maximum": 5}        },        "required": ["location"],        "additionalProperties": false      }    }  }]

Four choices in there are doing real work, and each one exists because of a specific failure:

  • "RIGHT NOW" versus "NEXT 1-5 days" in the descriptions. Without an explicit contrast the model calls get_current_weather for "what's it like tomorrow" perhaps one time in five, and then confidently reports today's temperature as tomorrow's.
  • enum on units. A free-text string gets you "metric", "celsius", "C", and "Metric" across four calls, and your handler has to guess.
  • maximum: 5 on days. The model does not know your upstream stops at five days unless the schema says so — and it will cheerfully ask for 10.
  • additionalProperties: false. Without it, an invented argument like "detailed": true lands in your handler's keyword arguments and raises TypeError: get_forecast() got an unexpected keyword argument 'detailed' in the middle of the loop.

A tool description is not documentation for humans who might read it later. It is the entire evidence the model has when deciding whether this tool answers the question in front of it.

Writing the handlers

Every handler obeys the same three rules: it validates what it was given, it never raises, and it returns the smallest useful result rather than the upstream's raw body.

Python
import time, requestsGEO_URL  = "https://api.openweathermap.org/geo/1.0/direct"BASE_URL = "https://api.openweathermap.org/data/2.5"_cache: dict[tuple, tuple[float, dict]] = {}      # key -> (expires_at, payload)def _cached(key: tuple, ttl_s: int, produce):    now = time.monotonic()    hit = _cache.get(key)    if hit and hit[0] > now:        return hit[1]    value = produce()    if value.get("ok"):                            # never cache a failure        _cache[key] = (now + ttl_s, value)    return valuedef _geocode(location: str) -> dict:    def fetch():        r = requests.get(GEO_URL, timeout=(3, 5), params={            "q": location, "limit": 1, "appid": WEATHER_API_KEY})        if r.status_code != 200:            return {"ok": False, "error": "geocode_failed"}        rows = r.json()        if not rows:            return {"ok": False, "error": "location_not_found",                    "message": f"No place matched {location!r}. "                               f"Try 'City, CC', e.g. 'Cambridge, GB'."}        return {"ok": True, "lat": rows[0]["lat"], "lon": rows[0]["lon"],                "label": f"{rows[0]['name']}, {rows[0].get('country', '')}"}    return _cached(("geo", location.lower()), 86_400, fetch)   # places don't movedef get_current_weather(location: str, units: str = "metric") -> dict:    if not isinstance(location, str) or not location.strip():        return {"ok": False, "error": "invalid_location"}    place = _geocode(location)    if not place["ok"]:        return place    def fetch():        try:            r = requests.get(f"{BASE_URL}/weather", timeout=(3, 5), params={                "lat": place["lat"], "lon": place["lon"],                "units": units, "appid": WEATHER_API_KEY})        except requests.RequestException:            return {"ok": False, "error": "weather_service_unreachable"}        if r.status_code != 200:            return {"ok": False, "error": "weather_service_error",                    "message": f"HTTP {r.status_code}"}        d = r.json()        return {"ok": True, "location": place["label"],                "temperature": round(d["main"]["temp"], 1),                "feels_like": round(d["main"]["feels_like"], 1),                "condition": d["weather"][0]["description"],                "humidity_pct": d["main"]["humidity"],                "units": "C" if units == "metric" else "F"}    return _cached(("now", place["label"], units), 600, fetch)   # 10-minute TTL

The result the model actually receives is small on purpose:

JSON
{  "ok": true,  "location": "Tokyo, JP",  "temperature": 22.4,  "feels_like": 21.1,  "condition": "broken clouds",  "humidity_pct": 65,  "units": "C"}

Note "units": "C" travelling back with the numbers. Without it the model has a bare 22.4 and has to infer the scale from context, which it will occasionally get wrong and report as 22 degrees Fahrenheit. Any number that has a unit should carry its unit.

What the 10-minute cache is worth

Weather does not change meaningfully in ten minutes, so a short time-to-live is free accuracy-wise and dramatic cost-wise. Take a realistic burst: 300 questions arrive in ten minutes, spread across 12 cities. Without the cache that is 300 geocode calls plus 300 weather calls. With it, 12 weather calls (one per city per window) and roughly zero geocode calls after the first day, since the coordinates of Leeds are cached for 24 hours. That is a 96 per cent reduction in upstream traffic, and it takes the free tier's 60-calls-per-minute limit from "constantly breached" to "never approached".

The if value.get("ok") guard matters as much as the cache itself. Caching a failure means one transient upstream blip becomes ten minutes of confidently reported failure for every user asking about that city.

The conversation loop

The loop uses the Chat Completions message format from the function-calling lesson; on the Responses API the same invariant holds with function_call items and function_call_output entries paired by call_id. The loop is a small state machine over a growing list of messages. Its one hard rule: every assistant message containing tool_calls must be followed by exactly one tool message per call, each carrying the matching tool_call_id. Violate it and you get the 400 that opened this lesson.

Python
import jsonimport jsonschemafrom openai import OpenAIclient = OpenAI(api_key=OPENAI_API_KEY)HANDLERS = {"get_current_weather": get_current_weather, "get_forecast": get_forecast}SYSTEM = ("You are a weather assistant. Use the tools for anything about actual "          "conditions; never guess a temperature. If a tool returns ok:false, "          "explain the problem plainly and suggest what the user could try.")class WeatherBot:    def __init__(self, max_tool_rounds: int = 4):        self.messages = [{"role": "system", "content": SYSTEM}]        self.max_tool_rounds = max_tool_rounds    def ask(self, user_input: str) -> str:        self.messages.append({"role": "user", "content": user_input})        for _ in range(self.max_tool_rounds):            reply = client.chat.completions.create(                model=MODEL, messages=self.messages,                tools=TOOLS, tool_choice="auto",            ).choices[0].message            # 1. ALWAYS append the assistant message, tool calls and all            self.messages.append(reply.model_dump(exclude_none=True))            if not reply.tool_calls:                return reply.content            # 2. One tool message per tool call, in the same order            for call in reply.tool_calls:                result = self._run(call.function.name, call.function.arguments)                self.messages.append({                    "role": "tool",                    "tool_call_id": call.id,            # the bit people forget                    "content": json.dumps(result),                })        return ("I got stuck looking that up after several attempts. "                "Could you rephrase the question?")    def _run(self, name: str, raw_args: str) -> dict:        handler = HANDLERS.get(name)        if handler is None:            return {"ok": False, "error": "unknown_tool", "message": f"No tool {name}."}        try:            args = json.loads(raw_args)        except json.JSONDecodeError:            return {"ok": False, "error": "unparseable_arguments",                    "message": "Arguments were not valid JSON. Try again."}        try:            jsonschema.validate(args, SCHEMAS[name])        except jsonschema.ValidationError as e:            return {"ok": False, "error": "invalid_arguments", "message": e.message}        try:            return handler(**args)        except Exception as e:                          # last line of defence            log.exception("handler %s failed", name)            return {"ok": False, "error": "tool_failed", "message": type(e).__name__}

Why the loop is capped

max_tool_rounds=4 is not decoration. Models can get into a rut — call a tool, read a confusing result, call the same tool again with a trivially different argument, forever. Every round is a full API call whose input is the entire conversation so far, so the cost climbs as the history grows.

Suppose the first request is 700 tokens of input and each round adds roughly 250 tokens of tool call plus tool result. A 12-round loop sends 700, 950, 1,200, … 3,450 tokens — an arithmetic series summing to 12 × (700 + 3450) / 2 = 24,900 input tokens. At 3 dollars per million that is 7.5 cents, against 0.45 cents for a normal single-tool turn: roughly 17 times the cost, for a question that never gets an answer. The cap converts an unbounded failure into a polite one.

What the message array looks like after one question

This is the single most useful thing to print while debugging. After "What's the weather in Tokyo?", self.messages holds:

JSON
[  {"role": "system", "content": "You are a weather assistant..."},  {"role": "user", "content": "What's the weather in Tokyo?"},  {"role": "assistant", "content": null,   "tool_calls": [     {"id": "call_9QkR2mVb8xNfLp", "type": "function",      "function": {"name": "get_current_weather",                   "arguments": "{\"location\":\"Tokyo\",\"units\":\"metric\"}"}}   ]},  {"role": "tool", "tool_call_id": "call_9QkR2mVb8xNfLp",   "content": "{\"ok\":true,\"location\":\"Tokyo, JP\",\"temperature\":22.4,\"condition\":\"broken clouds\",\"humidity_pct\":65,\"units\":\"C\"}"},  {"role": "assistant",   "content": "It's 22.4C in Tokyo with broken clouds and 65% humidity."}]

Two details are worth staring at. The arguments field is a JSON string, not an object — hence json.loads before you can touch it, and hence the possibility, however rare with modern constrained decoding, of it not parsing at all. And the content of a tool message is also a string; whatever your handler returned must be serialised. Return a datetime or a Mongo ObjectId from a handler and json.dumps raises TypeError: Object of type datetime is not JSON serializable, killing the turn.

Parallel tool calls

Ask "compare London and Cairo" and one assistant message may contain two tool calls at once. The loop above already handles this — it iterates reply.tool_calls — but only because it appends one tool message per call. Code written for a single call, using reply.tool_calls[0], produces the orphaned-call 400 the moment a user asks a comparison question.

Append the assistant message before you run anything, and one tool message per tool call afterwards. Almost every mysterious 400 from a function-calling loop is a violation of that pairing.

When the model gets the arguments wrong

The schema says days is between 1 and 5. A user asks "what's the weather doing over the next week?" and the model, reasonably enough, emits:

JSON
{"name": "get_forecast", "arguments": "{\"location\":\"Bristol\",\"days\":7}"}

Validation catches it before your handler runs, and the interesting part is what you do next. The wrong move is to raise, or to return "error". The right move is to hand the model a message it can act on:

JSON
{"ok": false, "error": "invalid_arguments", "message": "7 is greater than the maximum of 5"}

Feed that back as the tool result and the loop's next round produces:

JSON
{"name": "get_forecast", "arguments": "{\"location\":\"Bristol\",\"days\":5}"}

…followed by an answer that says "here are the next five days — that's as far ahead as I can see." The user gets a good answer, nobody sees a stack trace, and the self-correction cost one extra round trip. This is the payoff for making handlers total functions that return errors instead of raising them: the model is a retry mechanism, if you talk to it.

The same pattern covers a genuinely malformed response. Occasionally — far more often with older models or when the response is truncated by a low max_tokens — you get arguments like:

Text
{"location": "Bristol", "days":

json.loads raises JSONDecodeError: Expecting value: line 1 column 28 (char 27). Your _run catches it, returns unparseable_arguments, and the model reissues the call. Without that try/except the whole process falls over on a transient formatting glitch.

What arrivesCaught byReturned to the modelTypical outcome
Truncated JSONjson.loadsunparseable_argumentsModel reissues the call
days: 7jsonschemainvalid_arguments + reasonModel retries with 5
"detailed": trueadditionalProperties: falseinvalid_argumentsModel drops the field
City that does not existGeocode returning no rowslocation_not_found + hintModel asks the user to clarify
Upstream 503status_code != 200weather_service_errorModel explains, does not invent

Testing without touching the network

Tests that call the real weather API are slow, flaky, and quietly consume your quota; tests that call the real model are all of that plus non-deterministic. Split the test suite by what it actually exercises.

LevelWhat is fakedWhat it proves
Handler unit testsHTTP layerShaping, error mapping, caching, validation
Loop testsThe model clientMessage pairing, iteration cap, tool dispatch
Schema testsNothingBad arguments are rejected, good ones accepted
Live smoke testNothing (run manually)Keys work, upstream contract unchanged
Python
def test_unknown_city_returns_structured_error(monkeypatch):    monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(200, []))    r = get_current_weather("Nowhereville")    assert r["ok"] is False and r["error"] == "location_not_found"    assert "Try 'City, CC'" in r["message"]      # the hint the model needsdef test_failures_are_not_cached(monkeypatch):    monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(503, {}))    get_current_weather("Leeds")    monkeypatch.setattr(requests, "get", lambda *a, **k: FakeResponse(200, TOKYO_JSON))    assert get_current_weather("Leeds")["ok"] is True    # recovers immediatelydef test_days_above_maximum_is_rejected():    bot = WeatherBot()    r = bot._run("get_forecast", '{"location":"Bristol","days":7}')    assert r["error"] == "invalid_arguments" and "maximum of 5" in r["message"]def test_loop_stops_at_the_cap():    bot = WeatherBot(max_tool_rounds=3)    with always_calls_a_tool():                  # fake client, never finishes        answer = bot.ask("weather?")    assert "stuck" in answer    assert sum(1 for m in bot.messages if m["role"] == "tool") == 3def test_every_tool_call_has_a_matching_tool_message():    bot = WeatherBot()    with scripted_client([two_parallel_calls(), plain_answer()]):        bot.ask("compare London and Cairo")    ids = [c["id"] for m in bot.messages if m.get("tool_calls")                   for c in m["tool_calls"]]    responses = [m["tool_call_id"] for m in bot.messages if m["role"] == "tool"]    assert sorted(ids) == sorted(responses)

That last test is the one that earns its keep. It encodes the invariant behind the 400 at the top of this lesson, and it catches a regression the moment someone "simplifies" the loop to handle a single tool call.

Cost and latency, in numbers

A single tool-using turn makes two model calls, not one, plus the upstream fetch. Timed end to end on a cold cache:

StageTimeTokens
Model call 1 (decide + emit tool call)900 ms~700 in, ~40 out
Geocode (cache miss)180 ms—
Weather fetch350 ms—
Model call 2 (read result, write answer)1,400 ms~800 in, ~90 out
Total2,830 ms1,500 in, 130 out

At 3 dollars per million input tokens and 15 per million output, that is 1500/1e6 × 3 = 0.0045 plus 130/1e6 × 15 = 0.00195, so roughly 0.65 cents per question — about 6.50 dollars per thousand questions. On a warm cache the geocode disappears and the total drops to 2,650 ms.

The useful observation is where the time goes: 2,300 of those 2,830 milliseconds are the two model calls. Optimising your weather fetch from 350 ms to 200 ms is invisible to the user. Streaming the second model call so text appears as it is generated is not — it takes perceived latency from "nearly three seconds of nothing" to "about one second, then words".

When it breaks in front of a user

SymptomCauseFix
tool_call_ids did not have response messagesAssistant message not appended, or only the first tool call answeredAppend the assistant message first; loop over all tool_calls
TypeError: got an unexpected keyword argumentInvented argument reached handler(**args)additionalProperties: false plus schema validation before dispatch
Reports today's weather for "tomorrow"Tool descriptions do not contrast present with futureSay "RIGHT NOW" and "NEXT 1-5 days" explicitly
Invents a temperature, no tool call in the logtool_choice left unset, or the system prompt permits guessingtool_choice="auto" plus "never guess a temperature"
Same city, different answers minutes apartNo cache; upstream rounding differs per callCache successful lookups for 10 minutes
Slow, then a 401 an hour into a sessionos.getenv returned None for a missing keyos.environ["..."] so it fails at start-up
Conversation gets slower every turnHistory grows without bound; every turn resends it allTrim or summarise older turns above a token budget

Log the tool name, the arguments and the outcome for every call, tagged with a conversation id. Without that log, "it gave a weird answer yesterday" is unanswerable; with it, it is a two-minute lookup.

Where to take it

The interesting extensions are the ones that change the shape of the system rather than adding another endpoint:

  • Persistence. Store messages per user in SQLite or Redis so a session survives a restart. This immediately forces the history-trimming question, because a stored conversation grows forever and every turn pays for the whole thing.
  • A second, unrelated data source — air quality, or tide times. Two tools from one API is an easy choice for the model; four tools across two domains is where descriptions start to matter, and where you will see your first genuinely wrong tool selection.
  • An HTTP front end. Wrap ask() in a FastAPI endpoint keyed by session id. Now concurrency is real: the module-level _cache dictionary is shared across requests, which is fine for reads and wrong the moment you add anything stateful.
  • Streaming. The single biggest perceived-speed win available, worth more than any backend optimisation.
  • A tool that writes something — "text me if it rains tomorrow". The moment a tool has a side effect, everything changes: it needs idempotency so a retried loop does not send two texts, and it needs a per-conversation cap so a confused model cannot send forty.

Build the read-only version first and get the loop invariants right, because they do not change when the tools get more interesting. The bookkeeping you have just written — append the assistant message, answer every tool call by id, validate arguments, return errors instead of raising, cap the rounds, cache the reads, log everything — is the same bookkeeping behind an agent that manages calendars or files expenses. Only the handlers change.