Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Structured outputs you can parse
The PolicyPal web page needed more than text. It had to show each citation as a link to the right PDF page, show a "Talk to HR" button when the assistant was handing over, and hide the answer box entirely when a question was out of scope. The first version found these things with regular expressions over the reply. Within a week the model had written citations as "[1]", "(source 1)", "see Source 1" and "¹", and the regexes had grown to forty lines.
The problem was not the regexes. The problem was asking a machine to read prose meant for humans. When the application needs to act on a model's reply, the reply should be data with a fixed shape, checked like any other input.
This lesson turns PolicyPal's answer into a contract: a small schema, constrained generation to produce it, validation to check it, and a clear plan for the failures that remain.
Decide what the application needs
Start from the screen and the code, not from the model. Every field in the schema should exist because something reads it.
| Field | Type | Who uses it |
|---|---|---|
status | one of answered, not_in_policy, needs_human, out_of_scope | The UI decides which buttons to show; metrics count each status |
answer | text, at most 150 words | Shown to the user |
citations | list of source numbers | Turned into links to the PDF and page |
follow_up | text, empty if none | Shown as a suggested next question |
The status field is the most important design choice. It turns the model's judgement into a small, closed set that code can branch on. A common mistake is a confidence field with a number like 0.87. Models are poorly calibrated when asked to rate themselves, so that number looks precise and means little. A categorical status with clear definitions is more honest and easier to test.
The order matters too. status comes first, because in Section 7 PolicyPal streams the answer, and the page needs to know whether to show "Talk to HR" before the text starts arriving.
Constrained generation, then validation
Most providers can now constrain generation to a JSON schema, so the reply is guaranteed to parse and to have the required fields. In PolicyPal's client this is the schema argument, which the Anthropic adapter sends as output_config. Define the shape once in Pydantic and derive the schema from it.
1# policypal/schemas.py2from typing import Literal34from pydantic import BaseModel, ConfigDict, field_validator56Status = Literal["answered", "not_in_policy", "needs_human", "out_of_scope"]78class PolicyAnswer(BaseModel):9 model_config = ConfigDict(extra="forbid") # emits "additionalProperties": false1011 status: Status12 answer: str13 citations: list[int]14 follow_up: str # "" when there is nothing to suggest1516 @field_validator("answer")17 @classmethod18 def short_enough(cls, value: str) -> str:19 if len(value.split()) > 150:20 raise ValueError("answer is longer than 150 words")21 return value2223ANSWER_SCHEMA = PolicyAnswer.model_json_schema()Providers support a subset of JSON Schema for constrained generation. Types, enums, required fields and additionalProperties: false are safe everywhere. Length limits and number ranges are not always supported, so PolicyPal keeps them out of the schema and checks them in Pydantic instead. That is why short_enough is a validator and not a max_length on the field. Using "" instead of an optional None for follow_up keeps the schema simple for the same reason.
Valid JSON is not a correct answer
Constrained generation removes syntax errors. It does not remove three other failures, and your code must handle each.
- Truncation. If the reply hits
max_tokens, the JSON may be cut off. Check the stop reason before parsing. - Refusal. A provider's safety system may decline a request. The reply then has a refusal stop reason and may not match the schema at all.
- Valid but wrong. The JSON parses and the fields have the right types, yet the content breaks a business rule: citation 6 when only five sources were sent, or status
answeredwith no citations at all.
The third kind is the one teams forget, because the parser is happy. Only your code knows that citation 6 does not exist. This function handles all three.
1# policypal/answer.py2from pydantic import ValidationError34from policypal.schemas import ANSWER_SCHEMA, PolicyAnswer56FALLBACK = PolicyAnswer(status="needs_human", citations=[], follow_up="",7 answer="I couldn't produce a reliable answer. Please use Talk to HR.")89def problems(ans: PolicyAnswer, n_sources: int) -> list[str]:10 found = []11 if any(c < 1 or c > n_sources for c in ans.citations):12 found.append(f"citations must be between 1 and {n_sources}")13 if ans.status == "answered" and not ans.citations:14 found.append("an answered question must cite at least one source")15 return found1617def ask(llm, system: str, messages: list[dict], n_sources: int) -> PolicyAnswer:18 for attempt in range(2):19 reply = llm.complete(system, messages, schema=ANSWER_SCHEMA, max_tokens=600)20 if reply.stop_reason in ("max_tokens", "refusal"):21 return FALLBACK22 try:23 ans = PolicyAnswer.model_validate_json(reply.text)24 errors = problems(ans, n_sources)25 except ValidationError as exc:26 errors = [str(exc)]27 if not errors:28 return ans29 messages = messages + [30 {"role": "assistant", "content": reply.text},31 {"role": "user", "content": "Fix these problems and reply again: " + "; ".join(errors)},32 ]33 return FALLBACKThe loop gives the model exactly one chance to repair, with the specific error in plain words. A second failure returns a safe fallback that routes the person to a human. Truncation and refusal skip the repair, because repeating the same request is unlikely to help and costs money. The limit of one retry is deliberate: in the pilot, repairs that failed once failed again 80% of the time.
Measure the failure rates
Structure is only useful if you know how often each failure happens. Log the outcome of every call and look at the numbers weekly.
| Over 2,000 pilot questions | Free text plus regex | Structured output plus validation |
|---|---|---|
| Could not be parsed | 3.1% | 0% |
| Parsed but broke a business rule | not detected | 1.4% |
| Fixed by one repair | not possible | 1.1% |
| Fell back to "Talk to HR" | not possible | 0.3% |
The most interesting row is the second one. With free text, those 1.4% of replies were silently shown to users with broken or missing citations. Nobody knew they existed. Structure did not only make parsing easier. It made a hidden class of errors visible and countable.
When JSON is the wrong tool
Structured output is not free. JSON adds some tokens for keys and quotes, and it makes streaming harder, because a half-finished JSON object is not something you can show a user directly. For a pure chat reply with no application logic, plain text is fine. Use a schema when code needs to act on the reply: branch on a status, render links, store fields or call a tool. PolicyPal needs all four, so it uses a schema, and Section 7 shows how to stream the answer field as it arrives.
Check your understanding
0 of 3 answered
1.A reply parses perfectly and has status "answered", but its citations list is empty. What should PolicyPal do?
2.Why does PolicyPal use a status field with four fixed values instead of a confidence number from 0 to 1?
3.The reply stops with a max_tokens stop reason. Why does ask return the fallback immediately instead of trying a repair?