Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your pipeline expects strict JSON, but the LLM returns markdown-wrapped JSON or prose 3% of the time. Parsing fails. How do you guarantee structured output from an LLM?
What you need to know
Asking nicely ("respond only with JSON") works most of the time. At a million calls, "most of the time" is 30,000 failures. The fix is to make invalid output impossible at the point of generation.
The options, strongest first
| Method | Guarantees valid shape? | Where it's available |
|---|---|---|
| Provider structured-output mode with a JSON Schema | Yes, for supported schema features | Major hosted APIs, including OpenAI, Anthropic and Google |
| Grammar-constrained decoding | Yes | Self-hosted: vLLM structured outputs, Outlines, XGrammar, llama.cpp grammars |
| Tool or function calling with a schema | Usually; strict modes guarantee it | Most chat APIs |
| JSON mode (valid JSON, any shape) | Valid JSON only, not your schema | Many APIs |
| Prompt instructions plus tolerant parsing | No | Everywhere |
Define the schema once
1from typing import Literal2from pydantic import BaseModel, Field34class Extraction(BaseModel):5 invoice_no: str6 total: float = Field(description="Grand total including tax")7 currency: Literal["INR", "USD", "EUR"]8 po_number: str | None = Field(default=None, description="null if not printed")910result = llm.with_structured_output(Extraction).invoke(prompt) # LangChain; returns an ExtractionThe same Pydantic model produces the JSON Schema sent to the API and validates the result, so the two cannot drift apart. Literal limits the currency to three values; str | None gives the model a legal way to say "not present".
Defence in depth
Schema-valid is not the same as correct, and not every model supports constrained mode. So:
- Constrain — use structured-output mode wherever available.
- Validate — parse with Pydantic anyway; catch type or range errors the schema did not express.
- Repair once — on a validation error, send the error text back for one corrected attempt. A second failure goes to a review queue, not another retry.
- Tolerant fallback — for models without constrained mode: strip code fences, take the outermost balanced braces, then validate.
Design schemas that allow honesty
If a field is required but missing from the document, a constrained model must still fill it, so it invents something. Making fields optional, with null meaning "not present", prevents constrained decoding from forcing hallucinations.
The remaining failure: valid but wrong
Constrained decoding cannot stop "total": 1180.0 when the invoice says 11,800. That is a semantic error. A verifier that checks key values appear in the source text catches many of these.
Two separate metrics
- Parse-failure rate should drop to about zero.
- Field-level accuracy against a labelled set is the real quality metric.
Mixing them into one "success rate" hides which problem you have.
A real-life example
Scenario (illustrative numbers). A lending company extracts fields from salary slips for loan applications. Its prompt says "return only JSON", and about 3% of 200,000 monthly documents fail to parse because of fences or an "Here is the JSON:" preamble. Failed documents go to manual processing, which costs about ₹25 each.
The team switches to the provider's structured-output mode with a Pydantic schema and makes employer_pan and hra optional. Parse failures fall to zero, saving about ₹1.5 lakh a month in manual work. Field accuracy on a 500-slip labelled set is unchanged at 96%, which tells them the next project is accuracy, not parsing, and they add a verifier for net salary.
Follow-up questions to expect
- "Does constrained decoding hurt quality?" — It can slightly, if the schema forces an unnatural order or format; keep schemas simple and put reasoning fields before answer fields if you need them.
- "What schema features are unsupported?" — Providers support a subset of JSON Schema; very complex unions, recursion or some keywords may be rejected. Check the documentation and test.
- "Why not just retry until it parses?" — Retries cost money and latency and may never converge; constraint removes the failure instead of hoping it passes.