Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 3: Structured Output Integrity Breakdown


Scenario: a pure-Python pipeline expects JSON from the model, and every so often gets a fenced code block, a preamble sentence or a trailing explanation, and json.loads throws in production. How do you make it reliable?

What you need to know

json.loads fails because the model is a text generator: without a constraint, it sometimes adds the helpful sentence it was trained to write. The fix is to stop relying on the model's good behaviour.

Layer 1: make invalid output impossible

Structured-output modes constrain decoding to your JSON Schema. Define the schema once as a Pydantic model and derive everything from it:

Python
from pydantic import BaseModel, Field, ValidationErrorclass Invoice(BaseModel):    invoice_no: str    total: float = Field(ge=0)    gstin: str | None = Field(default=None, description="null if not printed")resp = client.responses.parse(model=MODEL, input=prompt, text_format=Invoice)  # OpenAI SDKinvoice = resp.output_parsed

The SDK turns Invoice into a strict JSON Schema for the request and parses the reply back into an Invoice. Other providers have equivalent schema-constrained modes; the Pydantic model stays the single source of truth.

Layer 2: validate, then repair once

Python
def parse_with_repair(raw: str) -> Invoice:    try:        return Invoice.model_validate_json(raw)    except ValidationError as e:        raw = call_model(f"{prompt}\n\nYour output failed validation:\n{e}\nReturn corrected JSON only.")        return Invoice.model_validate_json(raw)   # a second failure raises -> review queue

One retry, not a loop. If the model fails twice on the same input, the input is probably unusual (a blurred scan, a two-invoice PDF) and a person should look at it. A loop just burns tokens.

Layer 3: tolerant extraction as a backstop

For a model without constrained mode, clean before validating: strip Markdown code fences, then take the text from the first { to its matching closing brace. Always validate the result; never trust a cleanup step alone.

  1. Constrain — schema mode where the provider supports it.
  2. Validate — Pydantic on every response, even constrained ones.
  3. Repair once — send the error back; one attempt.
  4. Route — anything still invalid goes to a human queue with the raw output attached.

Schemas that permit honesty

If gstin is required and the invoice has none, a constrained model must produce a value, so it invents one. Optional fields with null for "not present" prevent that. This one choice often removes more bad data than any prompt change.

Two metrics, kept apart

MetricMeasuresTarget
Parse-failure rateDid we get schema-valid JSON?About zero
Field-level accuracyAre the values right?The real quality bar

Perfectly structured output can still contain a wrong total. Only the second metric tells you that.

A real-life example

Scenario (illustrative numbers). A chartered-accountancy firm's script extracts invoice data from 30,000 PDFs a month for GST filing. About 2% crash json.loads, and the nightly batch stops at the first crash, so staff arrive to a half-finished run twice a week.

The engineer adds structured-output mode with a Pydantic schema, one repair attempt, and a review queue instead of a crash. Parse failures drop to zero; 40 documents a month go to review, mostly scanned invoices with two pages merged. Making gstin optional also exposes a hidden problem: the old prompt had been inventing GSTINs for small suppliers who don't print one, which the field-level check now catches.

Follow-up questions to expect

  • "Why validate if the output is already constrained?" — Schemas can't express every rule (cross-field checks, formats like GSTIN), and fallback models may not be constrained.
  • "What goes in the repair prompt?" — The original prompt, the invalid output and the exact validation error; asking "try again" without the error rarely works.
  • "Should the pipeline stop on a failure?" — No. Isolate failures per document, continue the batch and report them.