Structured Output and Function Calling

Course Content

Why Structured Output Matters


A three-person team ships an expense extractor. The prompt is one sentence: "Extract the vendor, the amount and the date from this receipt text. Return JSON." They test it on fifty receipts. Fifty clean parses. They deploy it on a Thursday.

By Sunday there are 41 receipts sitting in a dead-letter queue and nobody knows why. The logs show json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0). Someone pulls the raw model output for one of them:

Text
Here's the extracted information:{"vendor": "Blue Bottle Coffee", "amount": 18.50, "date": "2026-03-14"}Let me know if you'd like me to pull out any other fields!

The JSON is perfect. The model just wrapped it in politeness. json.loads() hits the letter H at character zero and gives up. Nothing about the extraction was wrong — the packaging was wrong, and packaging is all a parser can see.

That is the whole subject in one incident. A language model produces text. Your code needs data. The gap between those two things is where a surprising fraction of production LLM bugs live, and closing it deliberately — rather than hoping — is what "structured output" means.

Triaging a ticket: prose against a contractThe free-text attempt• Answer arrives as a paragraph• Category worded a new way each run• Parsing is a regex you keep patching• Fails silently on the 51st receiptThe structured attempt• Answer arrives as named fields• Category drawn from a fixed enum• Parsing is one load, then validate• Failure is loud, at the boundary
Fifty clean parses prove nothing: free text has no contract, so the failure is deferred, not absent.

Why the model does this to you

It helps to be precise about what a language model is actually doing when you ask it for JSON. At each step it produces a probability distribution over the next token, and something samples from that distribution. Nothing in that machinery has a concept of "I opened a brace on line 1, so I owe a closing brace before I stop." There is no parser inside the model enforcing balance. There is only: given everything so far, what token tends to come next?

Your instruction "return JSON" is not a constraint on that process. It is evidence fed into it — text in the context that makes JSON-shaped continuations more probable. Strong evidence, usually. But "Here's the extracted information:" is also a highly probable way for an assistant to begin a reply, because assistants in the training data begin replies that way constantly. When the two pressures compete, the polite preamble sometimes wins.

An instruction in a prompt shifts probabilities. It does not create guarantees. Any design that treats "I asked it to" as "it will" is a design with an unhandled error path.

Three further properties of free-form generation make this worse than a one-off annoyance:

  • It is stochastic. The same input can produce differently packaged output on different calls. Even temperature=0 is not fully deterministic on hosted models, and many current reasoning models do not let you set a temperature at all. A test suite that passes fifty times proves less than it feels like it does.
  • It is input-sensitive. The failures cluster on unusual inputs — a receipt with two totals, a ticket written in three languages, a document that is mostly whitespace. Exactly the inputs your fifty-item test set did not contain.
  • It degrades quietly. A crash is the good outcome. The bad outcome is valid JSON with the wrong shape, which parses fine and poisons a database column three services downstream.

The catalogue of free-text failures

Before arguing that structure is worth the effort, it is worth seeing exactly how many distinct ways "just ask for JSON" breaks. These are all real shapes seen in production, from the same prompt.

FailureWhat the model emittedWhat your code does
PreambleSure! Here is the JSON: {...}Parse error at char 0
Markdown fenceTriple-backtick json block around the objectParse error at char 0
Trailing commentary{...} then "Note: the date was ambiguous."Parse error, "extra data"
Unquoted keys{vendor: "Blue Bottle", amount: 18.50}Parse error — that is JavaScript, not JSON
Single quotes{'vendor': 'Blue Bottle'}Parse error — that is a Python dict repr
Trailing comma{"a": 1, "b": 2,}Parse error at the closing brace
Truncation{"vendor": "Blue Bottle CoffParse error — hit the output token limit
Renamed field{"merchant": "Blue Bottle", ...}Parses. KeyError: 'vendor' later
Type drift{"amount": "18.50"} (string, not number)Parses. Sum of expenses becomes string concatenation
Invented category{"priority": "high-ish"}Parses. Router has no branch for it; ticket vanishes
Nesting drift{"result": {"vendor": ...}}Parses. Every field reads as None
Nulls for unknowns{"date": "unknown"} instead of nullParses. Date column rejects it, or stores garbage

Notice the split. The first seven are syntax failures: loud, immediate, easy to spot. The last five are shape and semantics failures: silent, delayed, and far more expensive. Any technique that only fixes the top half of that table has solved the cheap problem.

What "structured output" actually means

Structured output is a contract between the model and your code, and the contract has three separate layers. Confusing them is the single most common conceptual error in this area.

LayerThe promiseExample of a violation
SyntaxThe bytes parse as JSON{vendor: "x"
ShapeRequired fields present, types correct, no extras{"merchant": "x", "amount": "18.50"}
SemanticsValues are true and obey business rules{"amount": 1850.00} for an 18.50 receipt

No API feature gives you the third layer. Nothing forces a model to be right. What the techniques do is progressively guarantee the first two, so that your effort — and your review budget — goes where the real uncertainty is.

The spectrum of techniques

These are the tools available, roughly in order of increasing guarantee and increasing setup cost.

TechniqueGuaranteesCostUse when
Prompt only ("return JSON")NothingZeroPrototypes, one-off scripts
Prompt + few-shot examplesNothing, but noticeably better oddsExtra input tokensModels or providers with no JSON feature
Prompt + tolerant post-processingRecovers fenced or prefixed JSONFragile regex you will maintain foreverLegacy integrations you cannot change
JSON modeSyntax onlyOne API parameterShape varies or is genuinely open-ended
Schema-constrained generationSyntax and shapeWriting a schema; some feature limitsAnything that writes to a database
Function / tool callingSyntax and shape, plus which actionSchema per tool, plus an execution loopThe model must choose among capabilities
Client-side validationCatches everything the layer above missedA few lines; always worth itAlways — it is a net, not a replacement

The last row matters more than people expect. Validation on your side is cheap, provider-independent, and it is the only layer that survives a provider changing behaviour under you. Even with the strictest generation feature switched on, a truncated response or a safety refusal can still hand you something that is not your object.

Six reasons this is worth the effort

1. Integration with systems that were never going to read prose

A database column has a type. An HTTP endpoint has a request body schema. A message queue has a payload contract. A React component takes props. None of these will accept "The vendor appears to be Blue Bottle Coffee and the total was about eighteen fifty." The moment an LLM sits anywhere except directly in front of a human reader, its output has to become data, and the only question is whether that conversion is designed or improvised.

2. Reliability, measured

Put real numbers on it. Suppose you process 40,000 receipts a month and your prompt-only pipeline parses successfully 97.8% of the time — which is a realistic figure for a well-written prompt on a good model.

Text
Failures per month = 40,000 x (1 - 0.978)                   = 40,000 x 0.022                   = 880 receiptsHuman handling at 4 minutes each = 3,520 minutes                                  = 58.7 hours per month

Roughly 1.5 working weeks of somebody's month spent re-keying data because of packaging. That is the cost of skipping a schema, and it is a recurring cost, not a one-off.

3. Validation against business rules

Once output is a typed object rather than a paragraph, you can assert things about it before anything acts on it. An expense amount must be greater than zero and below the approval ceiling. A priority must be one of exactly four strings. A date must not be in the future. A currency code must be three uppercase letters.

You cannot write any of those checks against a sentence. You can write all of them against an object, in about six lines, and they will catch model mistakes that no amount of prompt engineering would have prevented.

4. Contract compliance with downstream APIs

When the model's output is forwarded to somebody else's API, their validator is now your error handler — and it will reject the whole request rather than the one bad field. If a payments API requires amount_cents as an integer and the model produced 18.50, you get a 400 with no partial credit. Producing the exact required shape at generation time removes an entire translation layer, and translation layers are where subtle bugs breed.

5. Cost and latency

Structured output is usually cheaper than prose, which surprises people who assume constraints cost extra. A conversational answer to "what is this receipt" runs around 210 output tokens once you count the greeting, the restatement and the offer to help. The equivalent JSON object runs around 48.

Text
Saving per call        = 210 - 48 = 162 output tokensSaving per month       = 40,000 x 162 = 6,480,000 tokensAt 15 dollars / 1M out = 6.48 x 15 = 97.20 dollars per monthLatency at 55 tokens/sec:  prose  210 / 55 = 3.82 s  JSON    48 / 55 = 0.87 s  saved                2.95 s per call

Three seconds per call is the difference between an interface that feels responsive and one that feels broken. And the saving compounds in an agent that makes six model calls to answer one question.

6. It is what makes automation safe enough to leave running

This is the reason that matters most, and it is a multiplication problem. Chain five steps together, each of which must parse and route correctly, at the 97.8% per-step rate from above:

Text
P(all five steps succeed) = 0.978 ^ 5                          = 0.8947                          -> 10.5% of runs derail somewhereWith strict schemas and validation, per-step 99.95%:P(all five succeed)       = 0.9995 ^ 5                          = 0.99750                          -> 1 run in 400

The same per-step reliability that is merely annoying in a single call is disqualifying in a chain. This is why structured output is not a tidiness preference — it is the precondition for building anything that runs more than one step without a human watching.

Per-step reliability compounds exponentially. An agent is only as good as its weakest parser, raised to the power of how many times it runs.

Worked example: triaging support tickets

Take a concrete task. Tickets arrive; each one must be routed to a queue, given a priority, and tagged. Here is the input:

Text
Subject: Charged twice for MarchI was billed 49.99 twice on March 3rd. I've been a customer forfour years and this is the second time. I need this refunded todayor I'm cancelling. Order #A-99213.

The free-text attempt

Prompt: "Analyse this support ticket." A representative response:

Text
This is a billing issue where the customer was double-charged49.99 on March 3rd. Given the repeat occurrence and the explicitcancellation threat, I'd treat this as fairly urgent — probablyhigh priority. The customer has been with you four years, soretention risk is real. The order reference is A-99213.

Everything in that paragraph is correct. It is also useless to a router. To act on it you would have to regex out "high priority", regex out "billing", regex out the order number — and every one of those regexes breaks the first time the model writes "I'd escalate this" instead of "high priority".

The structured attempt

Declare the contract first, as a schema:

JSON
{  "type": "object",  "properties": {    "category":    { "type": "string",                     "enum": ["billing", "technical", "account", "other"] },    "priority":    { "type": "string",                     "enum": ["low", "medium", "high", "urgent"] },    "order_id":    { "type": ["string", "null"],                     "pattern": "^A-[0-9]{5}$" },    "amount":      { "type": ["number", "null"], "minimum": 0 },    "churn_risk":  { "type": "boolean" },    "summary":     { "type": "string", "maxLength": 120 }  },  "required": ["category", "priority", "churn_risk", "summary"],  "additionalProperties": false}

And the conforming output:

JSON
{  "category": "billing",  "priority": "urgent",  "order_id": "A-99213",  "amount": 49.99,  "churn_risk": true,  "summary": "Duplicate charge of 49.99 on 3 March; repeat incident; threatens to cancel."}

Now routing is QUEUES[ticket["category"]]. Escalation is if ticket["churn_risk"] and ticket["priority"] == "urgent". No regex, no ambiguity, no rewrite when the model phrases things differently next Tuesday.

What a malformed response looks like, and how it is caught

Here is the one that actually hurts — because it parses:

JSON
{  "category": "billing/account",  "priority": "high-ish",  "order_id": "#A-99213",  "amount": "49.99",  "churn_risk": "yes",  "summary": "Customer double charged.",  "sentiment": "angry"}

json.loads() accepts this without complaint. Five separate violations are hiding in it: a category outside the enum, a priority outside the enum, an order_id with a stray # that fails the pattern, an amount that is a string rather than a number, a churn_risk that is a string rather than a boolean, and an extra sentiment field that additionalProperties: false forbids. Without validation, the router looks up QUEUES["billing/account"], gets a KeyError or a silent None, and the ticket disappears.

Validation turns all of that into one loud, specific error before anything acts:

Python
from jsonschema import Draft202012Validatorvalidator = Draft202012Validator(TICKET_SCHEMA)errors = sorted(validator.iter_errors(parsed), key=lambda e: e.path)for e in errors:    path = "/".join(str(p) for p in e.path) or "<root>"    print(f"{path}: {e.message}")# category: 'billing/account' is not one of ['billing','technical','account','other']# priority: 'high-ish' is not one of ['low','medium','high','urgent']# order_id: '#A-99213' does not match '^A-[0-9]{5}$'# amount: '49.99' is not of type 'number', 'null'# churn_risk: 'yes' is not of type 'boolean'# <root>: Additional properties are not allowed ('sentiment' was unexpected)

Six messages, each naming the exact field and the exact rule. That output is directly usable in three ways: as a log line for an on-call engineer, as the body of a retry prompt telling the model precisely what to fix, and as a metric you can chart to see which field drifts most often.

Misconceptions that cost people weeks

BeliefReality
"JSON mode means my schema is enforced."It enforces syntax only. Valid JSON with entirely the wrong fields is a conforming output.
"A strict schema means the values are correct."It means the values are the right type. A schema cannot tell you the model read 1850 instead of 18.50.
"Structured output makes the model dumber."The usual culprit is a bad schema, not the constraint. Cryptic field names, missing descriptions and impossible enums starve the model of context. Fix the schema before blaming the feature.
"I can just write a regex to pull out the JSON."You can, and it will work until the model emits two fenced blocks, or prose containing a brace, or a nested object your regex cannot balance. Regex cannot parse nested structures — this is a theoretical limit, not an effort problem.
"Validation is redundant once generation is constrained."Truncation at the token limit, safety refusals and network-level truncation all bypass the generation constraint. Validation is the only check that sees the bytes that actually arrived.
"Structure costs more tokens."Schemas cost input tokens once; prose costs output tokens on every call, and output tokens typically cost four to eight times as much as input tokens.

Designing the contract before you write the prompt

The practical shift is one of order. Most people write the prompt first and discover the data shape by looking at what came back. Do it the other way round.

Start from the consumer. Ask: what does the code on the other side actually need in order to act? For the ticket router that is a queue name, a priority, and a boolean for escalation. Not a sentiment score, not a confidence percentage, not a suggested reply — those are things somebody might want someday, and every optional field is another field the model can get wrong. A schema with four required fields fails less often than one with twelve, simply because there is less surface.

Then make the uncertain cases representable. If the model genuinely cannot tell which category a ticket belongs to, it needs somewhere to say so — an "other" enum value, a nullable field, a needs_review boolean. Give it no legitimate way to express doubt and it will express doubt illegitimately, by inventing "billing/account" or picking the nearest wrong option with total confidence. Most enum violations are the model trying to be honest in a schema that has no vocabulary for honesty.

Finally, decide what happens on a validation failure before you ship, because it will happen. The three sensible options are: retry once with the validator's error messages appended to the conversation, which fixes most transient drift; fall back to a queue for human review, which is right when the cost of a wrong action is high; or fail loudly and stop, which is right when the model output triggers something irreversible. What is never right is the fourth option — catching the exception, logging it, and carrying on with a half-populated object.

Design the object your code needs, then ask the model to fill it in. Reading a data shape out of whatever text came back is a habit that scales exactly as far as your first unusual input.