Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Structured outputs you can parse


The PolicyPal web page needed more than text. It had to show each citation as a link to the right PDF page, show a "Talk to HR" button when the assistant was handing over, and hide the answer box entirely when a question was out of scope. The first version found these things with regular expressions over the reply. Within a week the model had written citations as "[1]", "(source 1)", "see Source 1" and "¹", and the regexes had grown to forty lines.

The problem was not the regexes. The problem was asking a machine to read prose meant for humans. When the application needs to act on a model's reply, the reply should be data with a fixed shape, checked like any other input.

This lesson turns PolicyPal's answer into a contract: a small schema, constrained generation to produce it, validation to check it, and a clear plan for the failures that remain.

From model reply to trusted dataConstrained JSON from the schemaStop reason: truncated or refused?Pydantic: types and answer lengthBusiness rules: citations existRepair once, then hand to HR
Structure did not only make parsing easy; it exposed the 1.4% of replies that parsed fine but broke a rule.

Decide what the application needs

Start from the screen and the code, not from the model. Every field in the schema should exist because something reads it.

FieldTypeWho uses it
statusone of answered, not_in_policy, needs_human, out_of_scopeThe UI decides which buttons to show; metrics count each status
answertext, at most 150 wordsShown to the user
citationslist of source numbersTurned into links to the PDF and page
follow_uptext, empty if noneShown as a suggested next question

The status field is the most important design choice. It turns the model's judgement into a small, closed set that code can branch on. A common mistake is a confidence field with a number like 0.87. Models are poorly calibrated when asked to rate themselves, so that number looks precise and means little. A categorical status with clear definitions is more honest and easier to test.

The order matters too. status comes first, because in Section 7 PolicyPal streams the answer, and the page needs to know whether to show "Talk to HR" before the text starts arriving.

Constrained generation, then validation

Most providers can now constrain generation to a JSON schema, so the reply is guaranteed to parse and to have the required fields. In PolicyPal's client this is the schema argument, which the Anthropic adapter sends as output_config. Define the shape once in Pydantic and derive the schema from it.

Python
# policypal/schemas.pyfrom typing import Literalfrom pydantic import BaseModel, ConfigDict, field_validatorStatus = Literal["answered", "not_in_policy", "needs_human", "out_of_scope"]class PolicyAnswer(BaseModel):    model_config = ConfigDict(extra="forbid")   # emits "additionalProperties": false    status: Status    answer: str    citations: list[int]    follow_up: str                               # "" when there is nothing to suggest    @field_validator("answer")    @classmethod    def short_enough(cls, value: str) -> str:        if len(value.split()) > 150:            raise ValueError("answer is longer than 150 words")        return valueANSWER_SCHEMA = PolicyAnswer.model_json_schema()

Providers support a subset of JSON Schema for constrained generation. Types, enums, required fields and additionalProperties: false are safe everywhere. Length limits and number ranges are not always supported, so PolicyPal keeps them out of the schema and checks them in Pydantic instead. That is why short_enough is a validator and not a max_length on the field. Using "" instead of an optional None for follow_up keeps the schema simple for the same reason.

Valid JSON is not a correct answer

Constrained generation removes syntax errors. It does not remove three other failures, and your code must handle each.

  • Truncation. If the reply hits max_tokens, the JSON may be cut off. Check the stop reason before parsing.
  • Refusal. A provider's safety system may decline a request. The reply then has a refusal stop reason and may not match the schema at all.
  • Valid but wrong. The JSON parses and the fields have the right types, yet the content breaks a business rule: citation 6 when only five sources were sent, or status answered with no citations at all.

The third kind is the one teams forget, because the parser is happy. Only your code knows that citation 6 does not exist. This function handles all three.

Python
# policypal/answer.pyfrom pydantic import ValidationErrorfrom policypal.schemas import ANSWER_SCHEMA, PolicyAnswerFALLBACK = PolicyAnswer(status="needs_human", citations=[], follow_up="",                        answer="I couldn't produce a reliable answer. Please use Talk to HR.")def problems(ans: PolicyAnswer, n_sources: int) -> list[str]:    found = []    if any(c < 1 or c > n_sources for c in ans.citations):        found.append(f"citations must be between 1 and {n_sources}")    if ans.status == "answered" and not ans.citations:        found.append("an answered question must cite at least one source")    return founddef ask(llm, system: str, messages: list[dict], n_sources: int) -> PolicyAnswer:    for attempt in range(2):        reply = llm.complete(system, messages, schema=ANSWER_SCHEMA, max_tokens=600)        if reply.stop_reason in ("max_tokens", "refusal"):            return FALLBACK        try:            ans = PolicyAnswer.model_validate_json(reply.text)            errors = problems(ans, n_sources)        except ValidationError as exc:            errors = [str(exc)]        if not errors:            return ans        messages = messages + [            {"role": "assistant", "content": reply.text},            {"role": "user", "content": "Fix these problems and reply again: " + "; ".join(errors)},        ]    return FALLBACK

The loop gives the model exactly one chance to repair, with the specific error in plain words. A second failure returns a safe fallback that routes the person to a human. Truncation and refusal skip the repair, because repeating the same request is unlikely to help and costs money. The limit of one retry is deliberate: in the pilot, repairs that failed once failed again 80% of the time.

Measure the failure rates

Structure is only useful if you know how often each failure happens. Log the outcome of every call and look at the numbers weekly.

Over 2,000 pilot questionsFree text plus regexStructured output plus validation
Could not be parsed3.1%0%
Parsed but broke a business rulenot detected1.4%
Fixed by one repairnot possible1.1%
Fell back to "Talk to HR"not possible0.3%

The most interesting row is the second one. With free text, those 1.4% of replies were silently shown to users with broken or missing citations. Nobody knew they existed. Structure did not only make parsing easier. It made a hidden class of errors visible and countable.

When JSON is the wrong tool

Structured output is not free. JSON adds some tokens for keys and quotes, and it makes streaming harder, because a half-finished JSON object is not something you can show a user directly. For a pure chat reply with no application logic, plain text is fine. Use a schema when code needs to act on the reply: branch on a status, render links, store fields or call a tool. PolicyPal needs all four, so it uses a schema, and Section 7 shows how to stream the answer field as it arrives.

Check your understanding

0 of 3 answered

1.A reply parses perfectly and has status "answered", but its citations list is empty. What should PolicyPal do?

2.Why does PolicyPal use a status field with four fixed values instead of a confidence number from 0 to 1?

3.The reply stops with a max_tokens stop reason. Why does ask return the fallback immediately instead of trying a repair?