LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you ensure reliable structured outputs from LLMs?


What you need to know

Where it is available (2026)

  • Hosted APIs — JSON-schema structured output modes (for example response_format with a strict JSON schema on OpenAI-style APIs; similar features on other major providers). SDKs can take a Pydantic model directly.
  • vLLM — supports the OpenAI-style response_format, and its own structured_outputs request field ({"json": schema}, {"regex": ...}, {"choice": [...]}), backed by grammar engines such as xgrammar. The older guided_json field was removed in v0.12, so update old code.
  • SGLang, llama.cpp, Ollama — JSON schema or grammar support (Ollama through its format field).

The three layers of validation

  1. Syntax — guaranteed by constrained decoding.
  2. Schema — types, enums, patterns, ranges; checked with Pydantic.
  3. Business rules — does the transaction ID exist? Is the amount the same as the ledger? Only your code knows.
Python
from typing import Literalfrom pydantic import BaseModel, Field, ValidationErrorclass Dispute(BaseModel):    category: Literal["duplicate_charge", "fraud", "merchant_dispute", "other"]    txn_id: str = Field(pattern=r"^TXN[0-9]{10}$")    amount_inr: float = Field(gt=0, le=200_000)def validate(raw: str, known_txns: set[str]) -> Dispute:    d = Dispute.model_validate_json(raw)          # syntax + schema    if d.txn_id not in known_txns:                # business rule        raise ValueError(f"unknown transaction {d.txn_id}")    return dknown = {"TXN0000412345"}bad = '{"category": "duplicate_charge", "txn_id": "TXN0000499999", "amount_inr": 1499}'try:    validate(bad, known)except (ValidationError, ValueError) as e:    print("rejected:", e)     # rejected: unknown transaction TXN0000499999

The bad example is perfect JSON and matches the schema, but the transaction does not exist: the model invented it. Only the business check catches it. Libraries such as instructor automate "validate, send the error back, retry once".

Schema design tips

  • Keep it flat and small; deep nesting and many optional unions hurt quality.
  • Field names and descriptions are part of the prompt — refund_amount_inr is clearer than amt.
  • Put a reasoning or notes field before the answer fields if you want the model to think first.
  • Allow an explicit "unknown" or null so the model is not forced to invent.

A real-life example

A fintech support bot turns free-text complaints into dispute tickets. Before constrained decoding, about 3% of outputs failed to parse (markdown fences, trailing commas), and each failure became a human task. With a strict schema, parse failures fall to zero.

But the team's weekly review finds 1.2% of tickets with a well-formed transaction ID that does not exist — the model copied digits wrongly from the user's message. They add the ledger lookup; on a miss, the bot asks the user to pick from their last five transactions. Wrong-ticket rate falls to near zero, and the lesson goes into the runbook: schema-valid is not correct.

Follow-up questions to expect

  • "Does constrained decoding hurt quality?" — It can slightly, if the schema forces an unnatural format; flat schemas and a reasoning field first reduce this. Measure on your eval set.
  • "What about latency?" — The grammar is compiled once per new schema and cached, so the first request with a new schema is a little slower.
  • "What if you cannot use constrained decoding?" — Tool or function calling, a strict prompt with examples, then parse-validate-retry once.