Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Structured output and validating it
In the pilot, before structured output was switched on, about 1 reply in 200 was not valid JSON. Some began with "Here is the draft:". Some ended mid-object when the reply hit its token limit. The code called json.loads, got an exception, and the ticket showed an error to the agent.
Those were the easy failures, because they were loud. The quiet ones were worse: about 1 reply in 40 was perfectly valid JSON that was wrong. "Roti" instead of "Butter Roti", so the item did not match the order. "action": "refund" together with an escalation code. A refund of ₹180 on a line where the customer paid ₹150.
A model's output is untrusted input to your system, exactly like a form submitted from a browser. It needs validation before it reaches a screen, and certainly before it reaches money.
Three layers of checking
- Shape — is it JSON with the right fields and allowed values? The provider's structured-output feature handles most of this.
- Types and internal consistency — are numbers in range, do fields agree with each other? A Pydantic model handles this.
- Business sense — does the draft make sense against this order? Only your code knows the order, so only your code can check this.
Each layer catches things the one before cannot. A schema can say units is an integer, but not that it must be at most the quantity ordered. Pydantic can check that refund_inr equals the sum of the lines, but not that "Butter Roti" is in this particular order.
Layer 1: ask for the shape
Most providers now let you pass a JSON schema and will constrain the reply to it. Keep the schema simple: types, enums, required fields and additionalProperties: false. Put ranges and cross-field rules in your own validator, where you control the error messages and where they cannot be silently unsupported.
1NULLABLE_STR = {"anyOf": [{"type": "string"}, {"type": "null"}]}2CODES = ["food_safety", "high_value", "repeat_refunds", "needs_info", "out_of_scope", "unclear"]34DRAFT_SCHEMA = {5 "type": "object",6 "properties": {7 "action": {"type": "string", "enum": ["refund", "replace", "escalate"]},8 "lines": {"type": "array", "items": {9 "type": "object",10 "properties": {11 "item": {"type": "string"},12 "units": {"type": "integer"},13 "issue": {"type": "string", "enum": ["missing", "cold", "packaging", "wrong_item"]},14 "rule": {"type": "string", "enum": ["R1", "R2", "R3"]},15 "amount_inr": {"type": "integer"},16 },17 "required": ["item", "units", "issue", "rule", "amount_inr"],18 "additionalProperties": False,19 }},20 "refund_inr": {"type": "integer"},21 "reason": {"type": "string"},22 "escalation_code": {"anyOf": [{"type": "string", "enum": CODES}, {"type": "null"}]},23 "question_for_customer": NULLABLE_STR,24 },25 "required": ["action", "lines", "refund_inr", "reason", "escalation_code", "question_for_customer"],26 "additionalProperties": False,27}With this schema passed to llm.complete, syntax errors and stray text disappear. That is a real gain. But the schema says nothing about whether the draft is right.
Layer 2: types and internal consistency
1from typing import Literal2from pydantic import BaseModel, Field, model_validator34class Line(BaseModel):5 item: str6 units: int = Field(ge=1, le=50)7 issue: Literal["missing", "cold", "packaging", "wrong_item"]8 rule: Literal["R1", "R2", "R3"]9 amount_inr: int = Field(ge=0)1011class Draft(BaseModel):12 action: Literal["refund", "replace", "escalate"]13 lines: list[Line]14 refund_inr: int = Field(ge=0)15 reason: str = Field(min_length=10, max_length=400)16 escalation_code: Literal["food_safety", "high_value", "repeat_refunds",17 "needs_info", "out_of_scope", "unclear"] | None18 question_for_customer: str | None1920 @model_validator(mode="after")21 def consistent(self) -> "Draft":22 if self.action == "escalate":23 if self.escalation_code is None or self.lines:24 raise ValueError("escalate needs an escalation_code and no lines")25 elif self.escalation_code is not None or not self.lines:26 raise ValueError(f"{self.action} needs lines and escalation_code null")27 expected = sum(l.amount_inr for l in self.lines) if self.action == "refund" else 028 if self.refund_inr != expected:29 raise ValueError(f"refund_inr must be {expected} for action {self.action}")30 if self.escalation_code == "needs_info" and not self.question_for_customer:31 raise ValueError("needs_info requires question_for_customer")32 return selfField(ge=1, le=50) rejects zero or absurd unit counts. The Literal types repeat the schema's enums, so the model is checked even if a provider ignores part of the schema. The model_validator runs after the fields are parsed and checks that they agree: an escalation carries a code and no lines, a refund's total equals its lines, and a needs_info escalation carries the question for the customer.
The error messages are written for the model as much as for the logs. In a moment we send them back to it, so they must say what is wrong in plain words.
Layer 3: business sense against the order
1def check_against_order(draft: Draft, order: dict) -> list[str]:2 problems = []3 items = {item["name"]: item for item in order["items"]}4 for line in draft.lines:5 item = items.get(line.item)6 if item is None:7 problems.append(f"'{line.item}' is not in the order. Use one of: {', '.join(items)}")8 continue9 if line.units > item["qty"]:10 problems.append(f"{line.item}: {line.units} units affected, but only {item['qty']} ordered")11 if line.amount_inr > item["paid_inr"]:12 problems.append(f"{line.item}: {line.amount_inr} is more than the {item['paid_inr']} paid")13 return problemsThis function checks what only the order can tell you: that each item exists, that no more units are affected than were ordered, and that no line refunds more than was paid for it. It returns a list of problems rather than raising at the first one, so a single repair request can fix them all.
Notice what this does not do: it does not fuzzy-match "Roti" to "Butter Roti". Silent correction is tempting, but an order can contain "Butter Roti" and "Tandoori Roti", and guessing between them decides someone's money. Rejecting and asking again is safer, and the error message lists the valid names.
Repair once, then give up cleanly
1from pydantic import ValidationError2import llm34def draft_for_ticket(release, ticket_text: str, order: dict, minutes_late: int) -> Draft | None:5 base = render_input(ticket_text, order, minutes_late)6 user = base7 for attempt in range(2): # one try, plus one repair8 reply = llm.complete(release.system, user, model=release.model,9 schema=release.schema, max_tokens=release.max_tokens)10 try:11 draft = Draft.model_validate_json(reply.text)12 problems = check_against_order(draft, order)13 except ValidationError as err:14 problems = [e["msg"] for e in err.errors()]15 if not problems:16 return draft17 user = (f"{base}\n\nYour previous draft:\n{reply.text}\n\nIt was rejected because:\n- "18 + "\n- ".join(problems) + "\nReturn a corrected draft.")19 return None # shown as "red": the agent handles the ticket by handrelease is the versioned prompt from the next lesson; it carries the system prompt, schema and model id. The repair message includes the original input, the rejected draft and the exact problems, so the model does not have to guess what went wrong. Why only one repair? A draft that fails twice is usually a ticket the model cannot handle, and a third try rarely succeeds. Each attempt adds about 2 seconds and a full call's cost. The manual path already exists and is fine.
Log every rejection
Rejections are one of the most useful signals you have, and they cost nothing to collect. For each rejected draft, log the release version, which layer rejected it, the problem messages, and whether the repair succeeded.
Watch the rejection rate by layer over time. A jump in Pydantic rejections after a release usually means the prompt now confuses the model about the output. A jump in order-check rejections without any release often means the data changed: TiffinGo saw one the week restaurants added items like "Paneer Tikka (Half)" and "Paneer Tikka (Full)", and the model kept writing "Paneer Tikka". The fix was one sentence in the contract. The signal came from the rejection log two days before any agent complained.
Check your understanding
0 of 3 answered
1.The model returns "Roti" and the order contains "Butter Roti". What should the validator do?
2.Structured output guarantees the reply matches the JSON schema. Which check is still needed in your own code?
3.A draft fails validation twice. What should happen?