Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 3: Structured Output Integrity Breakdown
Scenario: a pure-Python pipeline expects JSON from the model, and every so often gets a fenced code block, a preamble sentence or a trailing explanation, and json.loads throws in production. How do you make it reliable?
What you need to know
json.loads fails because the model is a text generator: without a constraint, it sometimes adds the helpful sentence it was trained to write. The fix is to stop relying on the model's good behaviour.
Layer 1: make invalid output impossible
Structured-output modes constrain decoding to your JSON Schema. Define the schema once as a Pydantic model and derive everything from it:
1from pydantic import BaseModel, Field, ValidationError23class Invoice(BaseModel):4 invoice_no: str5 total: float = Field(ge=0)6 gstin: str | None = Field(default=None, description="null if not printed")78resp = client.responses.parse(model=MODEL, input=prompt, text_format=Invoice) # OpenAI SDK9invoice = resp.output_parsedThe SDK turns Invoice into a strict JSON Schema for the request and parses the reply back into an Invoice. Other providers have equivalent schema-constrained modes; the Pydantic model stays the single source of truth.
Layer 2: validate, then repair once
1def parse_with_repair(raw: str) -> Invoice:2 try:3 return Invoice.model_validate_json(raw)4 except ValidationError as e:5 raw = call_model(f"{prompt}\n\nYour output failed validation:\n{e}\nReturn corrected JSON only.")6 return Invoice.model_validate_json(raw) # a second failure raises -> review queueOne retry, not a loop. If the model fails twice on the same input, the input is probably unusual (a blurred scan, a two-invoice PDF) and a person should look at it. A loop just burns tokens.
Layer 3: tolerant extraction as a backstop
For a model without constrained mode, clean before validating: strip Markdown code fences, then take the text from the first { to its matching closing brace. Always validate the result; never trust a cleanup step alone.
- Constrain — schema mode where the provider supports it.
- Validate — Pydantic on every response, even constrained ones.
- Repair once — send the error back; one attempt.
- Route — anything still invalid goes to a human queue with the raw output attached.
Schemas that permit honesty
If gstin is required and the invoice has none, a constrained model must produce a value, so it invents one. Optional fields with null for "not present" prevent that. This one choice often removes more bad data than any prompt change.
Two metrics, kept apart
| Metric | Measures | Target |
|---|---|---|
| Parse-failure rate | Did we get schema-valid JSON? | About zero |
| Field-level accuracy | Are the values right? | The real quality bar |
Perfectly structured output can still contain a wrong total. Only the second metric tells you that.
A real-life example
Scenario (illustrative numbers). A chartered-accountancy firm's script extracts invoice data from 30,000 PDFs a month for GST filing. About 2% crash json.loads, and the nightly batch stops at the first crash, so staff arrive to a half-finished run twice a week.
The engineer adds structured-output mode with a Pydantic schema, one repair attempt, and a review queue instead of a crash. Parse failures drop to zero; 40 documents a month go to review, mostly scanned invoices with two pages merged. Making gstin optional also exposes a hidden problem: the old prompt had been inventing GSTINs for small suppliers who don't print one, which the field-level check now catches.
Follow-up questions to expect
- "Why validate if the output is already constrained?" — Schemas can't express every rule (cross-field checks, formats like GSTIN), and fallback models may not be constrained.
- "What goes in the repair prompt?" — The original prompt, the invalid output and the exact validation error; asking "try again" without the error rarely works.
- "Should the pipeline stop on a failure?" — No. Isolate failures per document, continue the batch and report them.