AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Structured output and validating it


In the pilot, before structured output was switched on, about 1 reply in 200 was not valid JSON. Some began with "Here is the draft:". Some ended mid-object when the reply hit its token limit. The code called json.loads, got an exception, and the ticket showed an error to the agent.

Those were the easy failures, because they were loud. The quiet ones were worse: about 1 reply in 40 was perfectly valid JSON that was wrong. "Roti" instead of "Butter Roti", so the item did not match the order. "action": "refund" together with an escalation code. A refund of ₹180 on a line where the customer paid ₹150.

A model's output is untrusted input to your system, exactly like a form submitted from a browser. It needs validation before it reaches a screen, and certainly before it reaches money.

Three layers, one repair, then the manual pathSchema —shape and enumsPydantic —ranges andconsistencyOrder check— items,units, paidOne repair withthe exact errorsStill failing:agent works it by handAbout 0.6% of pilot tickets ended on the manual path.
The schema removed every syntax error; only the order check could catch 'Roti' when the order says 'Butter Roti'.

Three layers of checking

  1. Shape — is it JSON with the right fields and allowed values? The provider's structured-output feature handles most of this.
  2. Types and internal consistency — are numbers in range, do fields agree with each other? A Pydantic model handles this.
  3. Business sense — does the draft make sense against this order? Only your code knows the order, so only your code can check this.

Each layer catches things the one before cannot. A schema can say units is an integer, but not that it must be at most the quantity ordered. Pydantic can check that refund_inr equals the sum of the lines, but not that "Butter Roti" is in this particular order.

Layer 1: ask for the shape

Most providers now let you pass a JSON schema and will constrain the reply to it. Keep the schema simple: types, enums, required fields and additionalProperties: false. Put ranges and cross-field rules in your own validator, where you control the error messages and where they cannot be silently unsupported.

Python
NULLABLE_STR = {"anyOf": [{"type": "string"}, {"type": "null"}]}CODES = ["food_safety", "high_value", "repeat_refunds", "needs_info", "out_of_scope", "unclear"]DRAFT_SCHEMA = {    "type": "object",    "properties": {        "action": {"type": "string", "enum": ["refund", "replace", "escalate"]},        "lines": {"type": "array", "items": {            "type": "object",            "properties": {                "item": {"type": "string"},                "units": {"type": "integer"},                "issue": {"type": "string", "enum": ["missing", "cold", "packaging", "wrong_item"]},                "rule": {"type": "string", "enum": ["R1", "R2", "R3"]},                "amount_inr": {"type": "integer"},            },            "required": ["item", "units", "issue", "rule", "amount_inr"],            "additionalProperties": False,        }},        "refund_inr": {"type": "integer"},        "reason": {"type": "string"},        "escalation_code": {"anyOf": [{"type": "string", "enum": CODES}, {"type": "null"}]},        "question_for_customer": NULLABLE_STR,    },    "required": ["action", "lines", "refund_inr", "reason", "escalation_code", "question_for_customer"],    "additionalProperties": False,}

With this schema passed to llm.complete, syntax errors and stray text disappear. That is a real gain. But the schema says nothing about whether the draft is right.

Layer 2: types and internal consistency

Python
from typing import Literalfrom pydantic import BaseModel, Field, model_validatorclass Line(BaseModel):    item: str    units: int = Field(ge=1, le=50)    issue: Literal["missing", "cold", "packaging", "wrong_item"]    rule: Literal["R1", "R2", "R3"]    amount_inr: int = Field(ge=0)class Draft(BaseModel):    action: Literal["refund", "replace", "escalate"]    lines: list[Line]    refund_inr: int = Field(ge=0)    reason: str = Field(min_length=10, max_length=400)    escalation_code: Literal["food_safety", "high_value", "repeat_refunds",                             "needs_info", "out_of_scope", "unclear"] | None    question_for_customer: str | None    @model_validator(mode="after")    def consistent(self) -> "Draft":        if self.action == "escalate":            if self.escalation_code is None or self.lines:                raise ValueError("escalate needs an escalation_code and no lines")        elif self.escalation_code is not None or not self.lines:            raise ValueError(f"{self.action} needs lines and escalation_code null")        expected = sum(l.amount_inr for l in self.lines) if self.action == "refund" else 0        if self.refund_inr != expected:            raise ValueError(f"refund_inr must be {expected} for action {self.action}")        if self.escalation_code == "needs_info" and not self.question_for_customer:            raise ValueError("needs_info requires question_for_customer")        return self

Field(ge=1, le=50) rejects zero or absurd unit counts. The Literal types repeat the schema's enums, so the model is checked even if a provider ignores part of the schema. The model_validator runs after the fields are parsed and checks that they agree: an escalation carries a code and no lines, a refund's total equals its lines, and a needs_info escalation carries the question for the customer.

The error messages are written for the model as much as for the logs. In a moment we send them back to it, so they must say what is wrong in plain words.

Layer 3: business sense against the order

Python
def check_against_order(draft: Draft, order: dict) -> list[str]:    problems = []    items = {item["name"]: item for item in order["items"]}    for line in draft.lines:        item = items.get(line.item)        if item is None:            problems.append(f"'{line.item}' is not in the order. Use one of: {', '.join(items)}")            continue        if line.units > item["qty"]:            problems.append(f"{line.item}: {line.units} units affected, but only {item['qty']} ordered")        if line.amount_inr > item["paid_inr"]:            problems.append(f"{line.item}: {line.amount_inr} is more than the {item['paid_inr']} paid")    return problems

This function checks what only the order can tell you: that each item exists, that no more units are affected than were ordered, and that no line refunds more than was paid for it. It returns a list of problems rather than raising at the first one, so a single repair request can fix them all.

Notice what this does not do: it does not fuzzy-match "Roti" to "Butter Roti". Silent correction is tempting, but an order can contain "Butter Roti" and "Tandoori Roti", and guessing between them decides someone's money. Rejecting and asking again is safer, and the error message lists the valid names.

Repair once, then give up cleanly

Python
from pydantic import ValidationErrorimport llmdef draft_for_ticket(release, ticket_text: str, order: dict, minutes_late: int) -> Draft | None:    base = render_input(ticket_text, order, minutes_late)    user = base    for attempt in range(2):  # one try, plus one repair        reply = llm.complete(release.system, user, model=release.model,                             schema=release.schema, max_tokens=release.max_tokens)        try:            draft = Draft.model_validate_json(reply.text)            problems = check_against_order(draft, order)        except ValidationError as err:            problems = [e["msg"] for e in err.errors()]        if not problems:            return draft        user = (f"{base}\n\nYour previous draft:\n{reply.text}\n\nIt was rejected because:\n- "                + "\n- ".join(problems) + "\nReturn a corrected draft.")    return None  # shown as "red": the agent handles the ticket by hand

release is the versioned prompt from the next lesson; it carries the system prompt, schema and model id. The repair message includes the original input, the rejected draft and the exact problems, so the model does not have to guess what went wrong. Why only one repair? A draft that fails twice is usually a ticket the model cannot handle, and a third try rarely succeeds. Each attempt adds about 2 seconds and a full call's cost. The manual path already exists and is fine.

Log every rejection

Rejections are one of the most useful signals you have, and they cost nothing to collect. For each rejected draft, log the release version, which layer rejected it, the problem messages, and whether the repair succeeded.

Watch the rejection rate by layer over time. A jump in Pydantic rejections after a release usually means the prompt now confuses the model about the output. A jump in order-check rejections without any release often means the data changed: TiffinGo saw one the week restaurants added items like "Paneer Tikka (Half)" and "Paneer Tikka (Full)", and the model kept writing "Paneer Tikka". The fix was one sentence in the contract. The signal came from the rejection log two days before any agent complained.

Check your understanding

0 of 3 answered

1.The model returns "Roti" and the order contains "Butter Roti". What should the validator do?

2.Structured output guarantees the reply matches the JSON schema. Which check is still needed in your own code?

3.A draft fails validation twice. What should happen?