Course Content
Structured Output and Function Calling
3 sections · 6 lessons
JSON Mode and Schema Validation
A team runs 12,000 product-description extractions a day. Someone reads the provider docs, finds JSON mode, adds one line to the request, and the parse errors stop dead. Zero JSONDecodeError for eleven days. The change gets written up in the sprint review as "JSON output now guaranteed".
On day twelve, the merchandising team reports that 340 products are showing a blank price band on the storefront. The pipeline logs are clean — no errors, no retries, no alerts. The stored objects look like this:
1{2 "title": "Merino Wool Crew Neck",3 "material": "merino wool",4 "price_band": "mid"5}and like this:
1{2 "product": {3 "title": "Merino Wool Crew Neck",4 "attributes": { "material": "merino wool", "pricing_tier": "mid" }5 }6}Both are valid JSON. JSON mode delivered exactly what it promised. It promised the wrong thing — or rather, the team assumed a promise that was never made. The second object has a different nesting, a renamed field, and a wrapper key, and every one of those is perfectly legal JSON.
The gap between "this parses" and "this is my object" is where this lesson lives. There are three distinct levels of guarantee available, they cost different amounts, and knowing precisely which one you have switched on is the difference between a pipeline you can trust and a pipeline that fails silently for eleven days.
Level 1: plain prompting, the baseline
The starting point is asking nicely and hoping.
1import os2from openai import OpenAI34client = OpenAI()5MODEL = os.environ.get("OPENAI_MODEL", "gpt-6-sol") # model ids live in config67response = client.responses.create(8 model=MODEL,9 input="Extract person info as JSON: Alice Johnson, Engineer, 5 years",10)11print(response.output_text)The examples in this lesson use OpenAI's Responses API, which OpenAI recommends for new projects. The older Chat Completions API is still supported and has the same three levels under different parameter names (response_format instead of text.format). Anthropic's equivalents appear later in the lesson.
Across repeated calls with identical input, that same request can return any of these:
{"name": "Alice Johnson", "role": "Engineer", "years": 5}Here's the JSON you asked for:```json{"name": "Alice Johnson", "role": "Engineer", "years": 5}```{name: Alice Johnson, role: Engineer, years: 5}{"person": {"full_name": "Alice Johnson", "job_title": "Engineer", "experience": "5 years"}}The first parses and matches. The second parses only after you strip a preamble and a fence. The third does not parse at all — unquoted keys are JavaScript object syntax, not JSON. The fourth parses, matches nothing, and will happily flow into your database as a row of nulls.
Prompt-only extraction is not useless. It is the right choice when you have no provider feature available, when the shape is genuinely open-ended, or when a human reads every result. It is the wrong choice for anything that runs unattended.
Level 2: JSON mode, a syntax guarantee
JSON mode is a generation setting, not a prompt technique. Turning it on changes what happens inside the decoding loop.
1import json23response = client.responses.create(4 model=MODEL,5 instructions="You output JSON objects only.",6 input="Extract person info as JSON: Alice Johnson, Engineer, 5 years",7 text={"format": {"type": "json_object"}},8)910result = json.loads(response.output_text)How the constraint actually works
This is worth understanding, because it explains precisely what JSON mode can and cannot do. At every generation step the model produces a score for every token in its vocabulary — perhaps 100,000 of them. Normally the sampler picks from that whole distribution. Under JSON mode, a small parser tracks the state of the JSON being built ("I am inside a string", "I have just closed a value and owe either a comma or a closing brace") and computes which tokens could legally come next. Every illegal token has its score set to negative infinity before sampling.
So the model literally cannot emit Here's at position zero, because H is not a legal first character of a JSON value. It cannot emit a trailing comma before }. It cannot leave a string unterminated by choosing an end-of-text token mid-string.
Constrained decoding does not persuade the model to behave. It removes the illegal options from the ballot before the vote is counted.
And that is exactly the boundary of what it does. The state machine knows JSON grammar. It knows nothing about your field names, so {"merchant": ...} and {"vendor": ...} are equally legal to it.
What JSON mode guarantees — and what it does not
| Guaranteed | Not guaranteed |
|---|---|
| Balanced braces and brackets | The fields you wanted are present |
| All keys and strings correctly quoted | Correct types (5 versus "5") |
| No trailing commas, no comments | Values drawn from an approved set |
| No prose before or after the object | Nesting matching your application's model |
| Correct escaping inside strings | A stable shape between two identical calls |
Three traps specific to JSON mode
The word "JSON" must appear in your input. OpenAI rejects a JSON-mode request whose instructions and input never mention JSON, with a 400. This is deliberate: it stops you enabling the flag on a chat endpoint and silently getting object-shaped answers to conversational questions. Put it in the instructions (the system message).
Truncation still breaks the output. This is the trap people trip over hardest, because it looks like the guarantee failing. If generation hits your max_tokens ceiling mid-object, the response is cut wherever it happened to be:
{"category": "billing", "priority": "urgent", "summary": "Customer was chThat is not valid JSON and json.loads() will raise on it. The constraint governs which token comes next; it has no power over a hard stop imposed from outside. Always check how the response ended. In OpenAI's Responses API a cut-off response has status == "incomplete" and incomplete_details.reason == "max_output_tokens" (Chat Completions reports finish_reason == "length"); Anthropic reports stop_reason == "max_tokens". If you truncated, no amount of retrying the parse will help. Raise the limit or shrink the requested object.
Empty or degenerate objects are legal. {} is valid JSON. So is {"result": "I could not determine the fields from this text."}. Both satisfy JSON mode completely.
Describing the shape: JSON Schema
To move from "any valid JSON" to "my object", you need a formal description of the shape. JSON Schema is the standard for that — a JSON document that describes other JSON documents.
1{2 "type": "object",3 "properties": {4 "name": {5 "type": "string",6 "description": "Full name of the person, as written in the source text"7 },8 "role": {9 "type": "string",10 "enum": ["Engineer", "Manager", "Designer"],11 "description": "Job title. Use the closest match from the list."12 },13 "experience_years": {14 "type": "integer",15 "minimum": 0,16 "maximum": 70,17 "description": "Whole years of professional experience"18 },19 "skills": {20 "type": "array",21 "items": { "type": "string" },22 "maxItems": 1023 }24 },25 "required": ["name", "role"],26 "additionalProperties": false27}Reading it keyword by keyword:
| Keyword | Meaning | Why it matters here |
|---|---|---|
type | The JSON type of this value | Stops "5" arriving where 5 was meant |
properties | The named fields of an object | Pins field names, killing merchant/vendor drift |
description | Free text about a field | The model reads this. It is prompt, not documentation |
enum | An exact allowed set of values | Turns a classification into a closed vocabulary |
minimum / maximum | Numeric bounds | Catches an age of 4,000 or a negative price |
pattern | A regex the string must match | Order IDs, currency codes, postcodes |
items | Schema for every array element | Stops a list of mixed types |
required | Fields that must be present | Anything not listed is optional and may vanish |
additionalProperties | Whether unlisted fields are allowed | Set false to reject invented fields |
The description row deserves emphasis. Those strings are sent to the model along with the schema. A field called amt with no description gets guessed at; the same field described as "Total charged to the customer in the transaction currency, excluding tax" gets extracted correctly far more often. Treat every description as a line of prompt that happens to live next to its field.
Level 3: schema-constrained generation
The third level applies the same masking trick as JSON mode, but with a state machine compiled from your schema rather than from generic JSON grammar. After the model emits {, the only legal next tokens are the ones beginning a key you declared. After "role":, the only legal continuations are the enum values you listed. A field can no longer be renamed, mistyped, omitted, or invented, because none of those paths exist in the machine.
1response = client.responses.create(2 model=MODEL,3 instructions="Extract structured person data.",4 input="Alice Johnson, Engineer, 5 years",5 text={6 "format": {7 "type": "json_schema",8 "name": "person",9 "strict": True,10 "schema": PERSON_SCHEMA, # must meet the rules in the next table11 }12 },13)14person = json.loads(response.output_text)The strict: True flag is what switches on the hard guarantee. Without it the schema is treated as a strong hint; with it, conformance is enforced by the decoder. One catch: the PERSON_SCHEMA above leaves experience_years and skills out of required, and OpenAI's strict mode rejects that. The restrictions below explain why, and the fix.
The restrictions strict mode imposes
Hard enforcement is not free. To compile a schema into a decoding automaton, the provider needs the schema to be finite and unambiguous, which rules out several things JSON Schema normally allows:
| Restriction | Consequence for you |
|---|---|
additionalProperties: false is mandatory on every object (OpenAI and Anthropic) | You must list every field you will accept |
OpenAI: every property must appear in required | There is no "optional field" — see the workaround below. Anthropic allows a property to be left out of required |
| Not every validation keyword is enforced, and support differs by provider | OpenAI's strict mode enforces pattern, format, minimum/maximum and minItems/maxItems. Anthropic's does not enforce numeric bounds, string lengths or pattern; its SDK moves them into the field description and checks them after the reply arrives |
| Schema size and shape are limited | Nesting depth and property counts are capped, and some providers reject recursive schemas; very deep trees must be flattened |
| The first call with a new schema is slower | The schema is compiled into a grammar and cached (Anthropic keeps it for 24 hours after last use); budget for a cold start |
OpenAI's "everything is required" rule catches people out constantly. The workaround, which works on every provider, is to make the type nullable rather than the field absent:
1{2 "order_id": {3 "type": ["string", "null"],4 "description": "Order reference if one is stated, otherwise null"5 }6}This is better design anyway. An absent field and a field that is explicitly null are different statements: "I did not answer" versus "I answered, and the answer is nothing". Forcing the model to say null makes the second case visible in your data instead of indistinguishable from a bug.
The row about unenforced keywords is the one that catches people silently. On a provider that does not enforce pattern, a schema saying ^A-[0-9]{5}$ still lets the model emit A-9921 with four digits, and the request succeeds without a warning. Support also changes over time, so check your provider's list of supported keywords rather than assuming. This is precisely why the client-side validation layer below is not optional.
Describing the schema with Pydantic instead
Writing raw JSON Schema by hand is tedious and easy to get subtly wrong. In Python, the usual approach is to declare the shape as a Pydantic model and let the library generate the schema.
1from typing import Literal, Optional2from pydantic import BaseModel, Field34class Ticket(BaseModel):5 category: Literal["billing", "technical", "account", "other"]6 priority: Literal["low", "medium", "high", "urgent"]7 order_id: Optional[str] = Field(8 None, description="Order reference like A-99213, or null if absent"9 )10 churn_risk: bool = Field(11 description="True if the customer threatens to leave"12 )13 summary: str = Field(max_length=120)1415# Generates the JSON Schema: enums, a nullable order_id, maxLength16schema = Ticket.model_json_schema()Two things come free. First, Literal[...] becomes an enum in the generated schema and a type your editor understands, so a typo in a downstream comparison is a red squiggle rather than a runtime surprise. Second, the same class validates the parsed response: Ticket.model_validate(parsed) either returns a typed object or raises with a field-by-field explanation. One declaration serves as prompt, contract and validator.
The OpenAI and Anthropic Python SDKs both have a helper that takes the model class directly and hands back a typed instance rather than a dict, so you never touch the schema JSON at all:
1from anthropic import Anthropic2from openai import OpenAI34# OpenAI, Responses API5r = OpenAI().responses.parse(model=MODEL, input=ticket_text, text_format=Ticket)6ticket = r.output_parsed # a Ticket instance78# Anthropic, Messages API9m = Anthropic().messages.parse(10 model=CLAUDE_MODEL, # from config, e.g. "claude-sonnet-5"11 max_tokens=1024,12 messages=[{"role": "user", "content": ticket_text}],13 output_format=Ticket,14)15ticket = m.parsed_outputUnder the hood each helper does the two steps above, plus one adjustment for its provider. The OpenAI helper sends a strict schema with every field in required, and the Anthropic helper moves constraints it cannot enforce, such as max_length=120, into the field description. Both then validate the reply against your full Pydantic model.
This is not one provider's feature
The mechanism is general. Anthropic's Messages API takes an output_config.format block to constrain the response shape and offers a messages.parse() helper that validates the result against your schema; on the tool side, setting strict: true on a tool definition (with additionalProperties: false and a complete required list) guarantees the tool input validates exactly. Google's Gemini API takes a response JSON schema in its generation config. Locally hosted open-weight models get the same behaviour from grammar-based samplers, which compile a schema or an EBNF grammar into the identical token mask. The parameter names differ; the automaton does not.
Client-side validation as a safety net
Even with strict generation enabled, validate what arrives. Four things get past the decoder:
- Truncation. A
max_tokenscut produces bytes that no constraint can rescue. - Refusals. A safety refusal is a different response shape entirely, not your object.
- Ignored keywords.
pattern,minimumandmaxLengthmay not constrain generation even when the provider accepts them in the schema. - Business rules no schema can express. "Refund amount must not exceed the original charge" is a cross-field rule; JSON Schema has no vocabulary for it.
1import json2from jsonschema import Draft202012Validator3from pydantic import ValidationError45validator = Draft202012Validator(TICKET_SCHEMA)67def parse_ticket(raw: str, truncated: bool):8 # truncated: OpenAI status == "incomplete", or Anthropic9 # stop_reason == "max_tokens"10 if truncated:11 return None, ["response truncated at the token limit"]12 try:13 data = json.loads(raw)14 except json.JSONDecodeError as e:15 return None, [f"not valid JSON: {e}"]1617 problems = [18 f"{'/'.join(str(p) for p in e.path) or 'root'}: {e.message}"19 for e in sorted(validator.iter_errors(data), key=lambda e: list(e.path))20 ]21 if problems:22 return None, problems2324 # Cross-field rules the schema cannot express25 if data["priority"] == "urgent" and not data["churn_risk"]:26 problems.append("urgent tickets must set churn_risk")2728 return (data if not problems else None), problemsReturning a list of readable problems rather than raising has a specific payoff: those strings are the ideal body for a repair prompt. Append the failed output and the error list to the conversation, ask for a corrected object, and most drift resolves on the second attempt.
The arithmetic of a retry loop
Take the 12,000-a-day pipeline. Measured validation failure rate on the current setup is 4.2%. A single retry carrying the error messages succeeds about 85% of the time.
Failures per day = 12,000 x 0.042 = 504Rescued by one retry = 504 x 0.85 = 428Still failing = 504 x 0.15 = 76Residual failure rate = 76 / 12,000 = 0.63%Extra API calls = 504 / 12,000 = +4.2% spendA 4.2% increase in call volume takes the unattended failure rate from 1 in 24 to 1 in 158. That is close to the best value-for-money available anywhere in this stack. But cap the loop at one or two attempts — a third attempt on the same input rescues almost nothing and turns a bad prompt into an expensive infinite loop.
Worked example: the ticket classifier at all three levels
Same input at each level, so the difference is visible:
Subject: Charged twice for MarchI was billed 49.99 twice on March 3rd. Second time this hashappened. Refund it today or I'm cancelling. Order A-99213.Level 1, prompt only. One representative response out of many possible:
Based on the ticket, here's my analysis in JSON:```json{"type": "Billing", "urgency": "High", "order": "A-99213"}```Fails to parse (preamble plus fence). Even after stripping both, the field names are type/urgency/order, not the ones your router reads, and the values are capitalised.
Level 2, JSON mode. Parses every time. But across a thousand calls you will see all of these:
{"category": "billing", "priority": "urgent", "order_id": "A-99213", "churn_risk": true, "summary": "Duplicate charge."}{"category": "Billing", "priority": "High", "order_id": "#A-99213", "churn_risk": "yes", "summary": "Duplicate charge.", "sentiment": "angry"}{"ticket": {"category": "billing", "priority": "urgent"}}All three are valid JSON. The second has three type or casing violations and an invented field; the third has moved everything one level down. Your router handles the first and mishandles the other two.
Level 3, strict schema. Every response is the first shape. Casing, nesting, extra fields and boolean-as-string are all unreachable states in the automaton. What remains uncertain is only whether the model judged the priority correctly — which is the part that actually needed a human's attention all along.
Comparing the three levels
| Prompt only | JSON mode | Strict schema | |
|---|---|---|---|
| Parses | Usually | Always (unless truncated) | Always (unless truncated) |
| Field names fixed | No | No | Yes |
| Types fixed | No | No | Yes |
| Enums respected | No | No | Yes |
| No extra fields | No | No | Yes |
| Values correct | No | No | No |
| Setup cost | None | One parameter | Write and maintain a schema |
| Model coverage | All | Most current models | Specific model versions |
| Right for | Prototypes, human-read output | Open-ended or varying shapes | Anything writing to a system |
JSON mode guarantees the envelope. A strict schema guarantees the form inside it. Neither guarantees the answer, and no future feature will.
Pitfalls and what to do about them
| Symptom | Cause | Fix |
|---|---|---|
| 400 error the moment JSON mode is enabled | The word "JSON" appears nowhere in the messages | Add it to the system prompt |
| Parse error despite JSON mode | Truncated at the output token limit | Check for truncation (status == "incomplete" on OpenAI, stop_reason == "max_tokens" on Anthropic); raise the limit or shrink the object |
| Schema rejected by the API in strict mode | Missing additionalProperties: false, or (on OpenAI) a property left out of required | Add both; make optional fields nullable instead |
| Enum value outside the list | Not using strict mode, or the input fits none of the options | Turn on strict mode and add an explicit "other" value |
| Numbers arriving as strings | Type not pinned, or the source text contains "49.99 USD" | Pin the type; put the unit in a separate field |
| Empty strings where null was meant | Non-nullable string type, so the model has no way to say "absent" | Use ["string", "null"] and say so in the description |
| Accuracy dropped after adding the schema | Cryptic field names, no descriptions, or an enum with no escape hatch | Rename fields to be self-explanatory, describe every one, add "other" |
pattern violated in strict mode | Your provider does not enforce that keyword | Validate client-side; do not rely on the generator for regex |
| Deeply nested output silently wrong | Schema exceeded the provider's depth or property cap | Flatten, or split into two calls |
| First request after a deploy is slow | Schema automaton compiling for the first time | Expected; keep the schema byte-stable so the cache holds |
Choosing a level, and living with it
The decision rule is short. If a human reads every output, prompt-only is fine and a schema is overhead. If code reads the output but the shape genuinely varies — say, extracting whatever fields happen to appear in an arbitrary document — JSON mode plus validation is the right trade, because a schema you cannot write in advance is not a schema. If code reads the output and the shape is known, use a strict schema, and treat any argument against it as an argument for spending your afternoons on parse errors instead.
Whichever level you pick, three habits make the difference in practice. Version the schema. Give it a name and a number, store it next to the code, and log which version produced each record — because the day you add a field is the day you need to tell old rows from new ones. Log every validation failure with the raw output attached. The error message alone tells you a field was wrong; the raw output tells you why, and after a hundred of them you will see a pattern worth fixing in the schema. Chart your failure rate as a metric. A model version change on the provider's side can move it from 0.6% to 4% overnight, and if the only place that shows up is a dead-letter queue nobody reads, you will find out from a customer.
Constrained generation moves the failure from your parser into your data quality report. That is a large improvement and it is not the same as the failure going away.