Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you ensure reliable structured outputs from LLMs?
What you need to know
Where it is available (2026)
- Hosted APIs — JSON-schema structured output modes (for example
response_formatwith a strict JSON schema on OpenAI-style APIs; similar features on other major providers). SDKs can take a Pydantic model directly. - vLLM — supports the OpenAI-style
response_format, and its ownstructured_outputsrequest field ({"json": schema},{"regex": ...},{"choice": [...]}), backed by grammar engines such as xgrammar. The olderguided_jsonfield was removed in v0.12, so update old code. - SGLang, llama.cpp, Ollama — JSON schema or grammar support (Ollama through its
formatfield).
The three layers of validation
- Syntax — guaranteed by constrained decoding.
- Schema — types, enums, patterns, ranges; checked with Pydantic.
- Business rules — does the transaction ID exist? Is the amount the same as the ledger? Only your code knows.
1from typing import Literal2from pydantic import BaseModel, Field, ValidationError34class Dispute(BaseModel):5 category: Literal["duplicate_charge", "fraud", "merchant_dispute", "other"]6 txn_id: str = Field(pattern=r"^TXN[0-9]{10}$")7 amount_inr: float = Field(gt=0, le=200_000)89def validate(raw: str, known_txns: set[str]) -> Dispute:10 d = Dispute.model_validate_json(raw) # syntax + schema11 if d.txn_id not in known_txns: # business rule12 raise ValueError(f"unknown transaction {d.txn_id}")13 return d1415known = {"TXN0000412345"}16bad = '{"category": "duplicate_charge", "txn_id": "TXN0000499999", "amount_inr": 1499}'17try:18 validate(bad, known)19except (ValidationError, ValueError) as e:20 print("rejected:", e) # rejected: unknown transaction TXN0000499999The bad example is perfect JSON and matches the schema, but the transaction does not exist: the model invented it. Only the business check catches it. Libraries such as instructor automate "validate, send the error back, retry once".
Schema design tips
- Keep it flat and small; deep nesting and many optional unions hurt quality.
- Field names and descriptions are part of the prompt —
refund_amount_inris clearer thanamt. - Put a
reasoningornotesfield before the answer fields if you want the model to think first. - Allow an explicit "unknown" or
nullso the model is not forced to invent.
A real-life example
A fintech support bot turns free-text complaints into dispute tickets. Before constrained decoding, about 3% of outputs failed to parse (markdown fences, trailing commas), and each failure became a human task. With a strict schema, parse failures fall to zero.
But the team's weekly review finds 1.2% of tickets with a well-formed transaction ID that does not exist — the model copied digits wrongly from the user's message. They add the ledger lookup; on a miss, the bot asks the user to pick from their last five transactions. Wrong-ticket rate falls to near zero, and the lesson goes into the runbook: schema-valid is not correct.
Follow-up questions to expect
- "Does constrained decoding hurt quality?" — It can slightly, if the schema forces an unnatural format; flat schemas and a reasoning field first reduce this. Measure on your eval set.
- "What about latency?" — The grammar is compiled once per new schema and cached, so the first request with a new schema is a little slower.
- "What if you cannot use constrained decoding?" — Tool or function calling, a strict prompt with examples, then parse-validate-retry once.