Structured Output and Function Calling

Course Content

JSON Mode and Schema Validation


A team runs 12,000 product-description extractions a day. Someone reads the provider docs, finds JSON mode, adds one line to the request, and the parse errors stop dead. Zero JSONDecodeError for eleven days. The change gets written up in the sprint review as "JSON output now guaranteed".

On day twelve, the merchandising team reports that 340 products are showing a blank price band on the storefront. The pipeline logs are clean — no errors, no retries, no alerts. The stored objects look like this:

JSON
{  "title": "Merino Wool Crew Neck",  "material": "merino wool",  "price_band": "mid"}

and like this:

JSON
{  "product": {    "title": "Merino Wool Crew Neck",    "attributes": { "material": "merino wool", "pricing_tier": "mid" }  }}

Both are valid JSON. JSON mode delivered exactly what it promised. It promised the wrong thing — or rather, the team assumed a promise that was never made. The second object has a different nesting, a renamed field, and a wrapper key, and every one of those is perfectly legal JSON.

The gap between "this parses" and "this is my object" is where this lesson lives. There are three distinct levels of guarantee available, they cost different amounts, and knowing precisely which one you have switched on is the difference between a pipeline you can trust and a pipeline that fails silently for eleven days.

Three levels of guarantee, plus a netPrompt only — hope it returns JSONJSON mode — syntax is guaranteed, shape is notSchema-constrained decoding — shape guaranteed tooClient-side validation — meaning and business rules
Each level removes one class of failure; none of them makes the values inside the fields true.

Level 1: plain prompting, the baseline

The starting point is asking nicely and hoping.

Python
import osfrom openai import OpenAIclient = OpenAI()MODEL = os.environ.get("OPENAI_MODEL", "gpt-6-sol")   # model ids live in configresponse = client.responses.create(    model=MODEL,    input="Extract person info as JSON: Alice Johnson, Engineer, 5 years",)print(response.output_text)

The examples in this lesson use OpenAI's Responses API, which OpenAI recommends for new projects. The older Chat Completions API is still supported and has the same three levels under different parameter names (response_format instead of text.format). Anthropic's equivalents appear later in the lesson.

Across repeated calls with identical input, that same request can return any of these:

Text
{"name": "Alice Johnson", "role": "Engineer", "years": 5}Here's the JSON you asked for:```json{"name": "Alice Johnson", "role": "Engineer", "years": 5}```{name: Alice Johnson, role: Engineer, years: 5}{"person": {"full_name": "Alice Johnson", "job_title": "Engineer",            "experience": "5 years"}}

The first parses and matches. The second parses only after you strip a preamble and a fence. The third does not parse at all — unquoted keys are JavaScript object syntax, not JSON. The fourth parses, matches nothing, and will happily flow into your database as a row of nulls.

Prompt-only extraction is not useless. It is the right choice when you have no provider feature available, when the shape is genuinely open-ended, or when a human reads every result. It is the wrong choice for anything that runs unattended.

Level 2: JSON mode, a syntax guarantee

JSON mode is a generation setting, not a prompt technique. Turning it on changes what happens inside the decoding loop.

Python
import jsonresponse = client.responses.create(    model=MODEL,    instructions="You output JSON objects only.",    input="Extract person info as JSON: Alice Johnson, Engineer, 5 years",    text={"format": {"type": "json_object"}},)result = json.loads(response.output_text)

How the constraint actually works

This is worth understanding, because it explains precisely what JSON mode can and cannot do. At every generation step the model produces a score for every token in its vocabulary — perhaps 100,000 of them. Normally the sampler picks from that whole distribution. Under JSON mode, a small parser tracks the state of the JSON being built ("I am inside a string", "I have just closed a value and owe either a comma or a closing brace") and computes which tokens could legally come next. Every illegal token has its score set to negative infinity before sampling.

So the model literally cannot emit Here's at position zero, because H is not a legal first character of a JSON value. It cannot emit a trailing comma before }. It cannot leave a string unterminated by choosing an end-of-text token mid-string.

Constrained decoding does not persuade the model to behave. It removes the illegal options from the ballot before the vote is counted.

And that is exactly the boundary of what it does. The state machine knows JSON grammar. It knows nothing about your field names, so {"merchant": ...} and {"vendor": ...} are equally legal to it.

What JSON mode guarantees — and what it does not

GuaranteedNot guaranteed
Balanced braces and bracketsThe fields you wanted are present
All keys and strings correctly quotedCorrect types (5 versus "5")
No trailing commas, no commentsValues drawn from an approved set
No prose before or after the objectNesting matching your application's model
Correct escaping inside stringsA stable shape between two identical calls

Three traps specific to JSON mode

The word "JSON" must appear in your input. OpenAI rejects a JSON-mode request whose instructions and input never mention JSON, with a 400. This is deliberate: it stops you enabling the flag on a chat endpoint and silently getting object-shaped answers to conversational questions. Put it in the instructions (the system message).

Truncation still breaks the output. This is the trap people trip over hardest, because it looks like the guarantee failing. If generation hits your max_tokens ceiling mid-object, the response is cut wherever it happened to be:

Text
{"category": "billing", "priority": "urgent", "summary": "Customer was ch

That is not valid JSON and json.loads() will raise on it. The constraint governs which token comes next; it has no power over a hard stop imposed from outside. Always check how the response ended. In OpenAI's Responses API a cut-off response has status == "incomplete" and incomplete_details.reason == "max_output_tokens" (Chat Completions reports finish_reason == "length"); Anthropic reports stop_reason == "max_tokens". If you truncated, no amount of retrying the parse will help. Raise the limit or shrink the requested object.

Empty or degenerate objects are legal. {} is valid JSON. So is {"result": "I could not determine the fields from this text."}. Both satisfy JSON mode completely.

Describing the shape: JSON Schema

To move from "any valid JSON" to "my object", you need a formal description of the shape. JSON Schema is the standard for that — a JSON document that describes other JSON documents.

JSON
{  "type": "object",  "properties": {    "name": {      "type": "string",      "description": "Full name of the person, as written in the source text"    },    "role": {      "type": "string",      "enum": ["Engineer", "Manager", "Designer"],      "description": "Job title. Use the closest match from the list."    },    "experience_years": {      "type": "integer",      "minimum": 0,      "maximum": 70,      "description": "Whole years of professional experience"    },    "skills": {      "type": "array",      "items": { "type": "string" },      "maxItems": 10    }  },  "required": ["name", "role"],  "additionalProperties": false}

Reading it keyword by keyword:

KeywordMeaningWhy it matters here
typeThe JSON type of this valueStops "5" arriving where 5 was meant
propertiesThe named fields of an objectPins field names, killing merchant/vendor drift
descriptionFree text about a fieldThe model reads this. It is prompt, not documentation
enumAn exact allowed set of valuesTurns a classification into a closed vocabulary
minimum / maximumNumeric boundsCatches an age of 4,000 or a negative price
patternA regex the string must matchOrder IDs, currency codes, postcodes
itemsSchema for every array elementStops a list of mixed types
requiredFields that must be presentAnything not listed is optional and may vanish
additionalPropertiesWhether unlisted fields are allowedSet false to reject invented fields

The description row deserves emphasis. Those strings are sent to the model along with the schema. A field called amt with no description gets guessed at; the same field described as "Total charged to the customer in the transaction currency, excluding tax" gets extracted correctly far more often. Treat every description as a line of prompt that happens to live next to its field.

Level 3: schema-constrained generation

The third level applies the same masking trick as JSON mode, but with a state machine compiled from your schema rather than from generic JSON grammar. After the model emits {, the only legal next tokens are the ones beginning a key you declared. After "role":, the only legal continuations are the enum values you listed. A field can no longer be renamed, mistyped, omitted, or invented, because none of those paths exist in the machine.

Python
response = client.responses.create(    model=MODEL,    instructions="Extract structured person data.",    input="Alice Johnson, Engineer, 5 years",    text={        "format": {            "type": "json_schema",            "name": "person",            "strict": True,            "schema": PERSON_SCHEMA,   # must meet the rules in the next table        }    },)person = json.loads(response.output_text)

The strict: True flag is what switches on the hard guarantee. Without it the schema is treated as a strong hint; with it, conformance is enforced by the decoder. One catch: the PERSON_SCHEMA above leaves experience_years and skills out of required, and OpenAI's strict mode rejects that. The restrictions below explain why, and the fix.

The restrictions strict mode imposes

Hard enforcement is not free. To compile a schema into a decoding automaton, the provider needs the schema to be finite and unambiguous, which rules out several things JSON Schema normally allows:

RestrictionConsequence for you
additionalProperties: false is mandatory on every object (OpenAI and Anthropic)You must list every field you will accept
OpenAI: every property must appear in requiredThere is no "optional field" — see the workaround below. Anthropic allows a property to be left out of required
Not every validation keyword is enforced, and support differs by providerOpenAI's strict mode enforces pattern, format, minimum/maximum and minItems/maxItems. Anthropic's does not enforce numeric bounds, string lengths or pattern; its SDK moves them into the field description and checks them after the reply arrives
Schema size and shape are limitedNesting depth and property counts are capped, and some providers reject recursive schemas; very deep trees must be flattened
The first call with a new schema is slowerThe schema is compiled into a grammar and cached (Anthropic keeps it for 24 hours after last use); budget for a cold start

OpenAI's "everything is required" rule catches people out constantly. The workaround, which works on every provider, is to make the type nullable rather than the field absent:

JSON
{  "order_id": {    "type": ["string", "null"],    "description": "Order reference if one is stated, otherwise null"  }}

This is better design anyway. An absent field and a field that is explicitly null are different statements: "I did not answer" versus "I answered, and the answer is nothing". Forcing the model to say null makes the second case visible in your data instead of indistinguishable from a bug.

The row about unenforced keywords is the one that catches people silently. On a provider that does not enforce pattern, a schema saying ^A-[0-9]{5}$ still lets the model emit A-9921 with four digits, and the request succeeds without a warning. Support also changes over time, so check your provider's list of supported keywords rather than assuming. This is precisely why the client-side validation layer below is not optional.

Describing the schema with Pydantic instead

Writing raw JSON Schema by hand is tedious and easy to get subtly wrong. In Python, the usual approach is to declare the shape as a Pydantic model and let the library generate the schema.

Python
from typing import Literal, Optionalfrom pydantic import BaseModel, Fieldclass Ticket(BaseModel):    category: Literal["billing", "technical", "account", "other"]    priority: Literal["low", "medium", "high", "urgent"]    order_id: Optional[str] = Field(        None, description="Order reference like A-99213, or null if absent"    )    churn_risk: bool = Field(        description="True if the customer threatens to leave"    )    summary: str = Field(max_length=120)# Generates the JSON Schema: enums, a nullable order_id, maxLengthschema = Ticket.model_json_schema()

Two things come free. First, Literal[...] becomes an enum in the generated schema and a type your editor understands, so a typo in a downstream comparison is a red squiggle rather than a runtime surprise. Second, the same class validates the parsed response: Ticket.model_validate(parsed) either returns a typed object or raises with a field-by-field explanation. One declaration serves as prompt, contract and validator.

The OpenAI and Anthropic Python SDKs both have a helper that takes the model class directly and hands back a typed instance rather than a dict, so you never touch the schema JSON at all:

Python
from anthropic import Anthropicfrom openai import OpenAI# OpenAI, Responses APIr = OpenAI().responses.parse(model=MODEL, input=ticket_text, text_format=Ticket)ticket = r.output_parsed                  # a Ticket instance# Anthropic, Messages APIm = Anthropic().messages.parse(    model=CLAUDE_MODEL,                   # from config, e.g. "claude-sonnet-5"    max_tokens=1024,    messages=[{"role": "user", "content": ticket_text}],    output_format=Ticket,)ticket = m.parsed_output

Under the hood each helper does the two steps above, plus one adjustment for its provider. The OpenAI helper sends a strict schema with every field in required, and the Anthropic helper moves constraints it cannot enforce, such as max_length=120, into the field description. Both then validate the reply against your full Pydantic model.

This is not one provider's feature

The mechanism is general. Anthropic's Messages API takes an output_config.format block to constrain the response shape and offers a messages.parse() helper that validates the result against your schema; on the tool side, setting strict: true on a tool definition (with additionalProperties: false and a complete required list) guarantees the tool input validates exactly. Google's Gemini API takes a response JSON schema in its generation config. Locally hosted open-weight models get the same behaviour from grammar-based samplers, which compile a schema or an EBNF grammar into the identical token mask. The parameter names differ; the automaton does not.

Client-side validation as a safety net

Even with strict generation enabled, validate what arrives. Four things get past the decoder:

  • Truncation. A max_tokens cut produces bytes that no constraint can rescue.
  • Refusals. A safety refusal is a different response shape entirely, not your object.
  • Ignored keywords. pattern, minimum and maxLength may not constrain generation even when the provider accepts them in the schema.
  • Business rules no schema can express. "Refund amount must not exceed the original charge" is a cross-field rule; JSON Schema has no vocabulary for it.
Python
import jsonfrom jsonschema import Draft202012Validatorfrom pydantic import ValidationErrorvalidator = Draft202012Validator(TICKET_SCHEMA)def parse_ticket(raw: str, truncated: bool):    # truncated: OpenAI status == "incomplete", or Anthropic    # stop_reason == "max_tokens"    if truncated:        return None, ["response truncated at the token limit"]    try:        data = json.loads(raw)    except json.JSONDecodeError as e:        return None, [f"not valid JSON: {e}"]    problems = [        f"{'/'.join(str(p) for p in e.path) or 'root'}: {e.message}"        for e in sorted(validator.iter_errors(data), key=lambda e: list(e.path))    ]    if problems:        return None, problems    # Cross-field rules the schema cannot express    if data["priority"] == "urgent" and not data["churn_risk"]:        problems.append("urgent tickets must set churn_risk")    return (data if not problems else None), problems

Returning a list of readable problems rather than raising has a specific payoff: those strings are the ideal body for a repair prompt. Append the failed output and the error list to the conversation, ask for a corrected object, and most drift resolves on the second attempt.

The arithmetic of a retry loop

Take the 12,000-a-day pipeline. Measured validation failure rate on the current setup is 4.2%. A single retry carrying the error messages succeeds about 85% of the time.

Text
Failures per day        = 12,000 x 0.042        = 504Rescued by one retry    =    504 x 0.85         = 428Still failing           =    504 x 0.15         =  76Residual failure rate   = 76 / 12,000           = 0.63%Extra API calls         = 504 / 12,000          = +4.2% spend

A 4.2% increase in call volume takes the unattended failure rate from 1 in 24 to 1 in 158. That is close to the best value-for-money available anywhere in this stack. But cap the loop at one or two attempts — a third attempt on the same input rescues almost nothing and turns a bad prompt into an expensive infinite loop.

Worked example: the ticket classifier at all three levels

Same input at each level, so the difference is visible:

Text
Subject: Charged twice for MarchI was billed 49.99 twice on March 3rd. Second time this hashappened. Refund it today or I'm cancelling. Order A-99213.

Level 1, prompt only. One representative response out of many possible:

Text
Based on the ticket, here's my analysis in JSON:```json{"type": "Billing", "urgency": "High", "order": "A-99213"}```

Fails to parse (preamble plus fence). Even after stripping both, the field names are type/urgency/order, not the ones your router reads, and the values are capitalised.

Level 2, JSON mode. Parses every time. But across a thousand calls you will see all of these:

JSON
{"category": "billing", "priority": "urgent", "order_id": "A-99213", "churn_risk": true, "summary": "Duplicate charge."}
JSON
{"category": "Billing", "priority": "High", "order_id": "#A-99213", "churn_risk": "yes", "summary": "Duplicate charge.", "sentiment": "angry"}
JSON
{"ticket": {"category": "billing", "priority": "urgent"}}

All three are valid JSON. The second has three type or casing violations and an invented field; the third has moved everything one level down. Your router handles the first and mishandles the other two.

Level 3, strict schema. Every response is the first shape. Casing, nesting, extra fields and boolean-as-string are all unreachable states in the automaton. What remains uncertain is only whether the model judged the priority correctly — which is the part that actually needed a human's attention all along.

Comparing the three levels

Prompt onlyJSON modeStrict schema
ParsesUsuallyAlways (unless truncated)Always (unless truncated)
Field names fixedNoNoYes
Types fixedNoNoYes
Enums respectedNoNoYes
No extra fieldsNoNoYes
Values correctNoNoNo
Setup costNoneOne parameterWrite and maintain a schema
Model coverageAllMost current modelsSpecific model versions
Right forPrototypes, human-read outputOpen-ended or varying shapesAnything writing to a system

JSON mode guarantees the envelope. A strict schema guarantees the form inside it. Neither guarantees the answer, and no future feature will.

Pitfalls and what to do about them

SymptomCauseFix
400 error the moment JSON mode is enabledThe word "JSON" appears nowhere in the messagesAdd it to the system prompt
Parse error despite JSON modeTruncated at the output token limitCheck for truncation (status == "incomplete" on OpenAI, stop_reason == "max_tokens" on Anthropic); raise the limit or shrink the object
Schema rejected by the API in strict modeMissing additionalProperties: false, or (on OpenAI) a property left out of requiredAdd both; make optional fields nullable instead
Enum value outside the listNot using strict mode, or the input fits none of the optionsTurn on strict mode and add an explicit "other" value
Numbers arriving as stringsType not pinned, or the source text contains "49.99 USD"Pin the type; put the unit in a separate field
Empty strings where null was meantNon-nullable string type, so the model has no way to say "absent"Use ["string", "null"] and say so in the description
Accuracy dropped after adding the schemaCryptic field names, no descriptions, or an enum with no escape hatchRename fields to be self-explanatory, describe every one, add "other"
pattern violated in strict modeYour provider does not enforce that keywordValidate client-side; do not rely on the generator for regex
Deeply nested output silently wrongSchema exceeded the provider's depth or property capFlatten, or split into two calls
First request after a deploy is slowSchema automaton compiling for the first timeExpected; keep the schema byte-stable so the cache holds

Choosing a level, and living with it

The decision rule is short. If a human reads every output, prompt-only is fine and a schema is overhead. If code reads the output but the shape genuinely varies — say, extracting whatever fields happen to appear in an arbitrary document — JSON mode plus validation is the right trade, because a schema you cannot write in advance is not a schema. If code reads the output and the shape is known, use a strict schema, and treat any argument against it as an argument for spending your afternoons on parse errors instead.

Whichever level you pick, three habits make the difference in practice. Version the schema. Give it a name and a number, store it next to the code, and log which version produced each record — because the day you add a field is the day you need to tell old rows from new ones. Log every validation failure with the raw output attached. The error message alone tells you a field was wrong; the raw output tells you why, and after a hundred of them you will see a pattern worth fixing in the schema. Chart your failure rate as a metric. A model version change on the provider's side can move it from 0.6% to 4% overnight, and if the only place that shows up is a dead-letter queue nobody reads, you will find out from a customer.

Constrained generation moves the failure from your parser into your data quality report. That is a large improvement and it is not the same as the failure going away.