Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 4: Structured Output Instability


Scenario: a LangChain chain feeds JSON into a downstream service, and occasionally the model returns prose or a fenced block and parsing fails. How do you make the output reliable?

What you need to know

An output parser on raw text is a repair step: the model writes whatever it likes, and you try to fix it. with_structured_output is a prevention step: the model is asked, through the provider's native mechanism, to produce exactly your schema.

The modern pattern

Python
from typing import Literalfrom pydantic import BaseModel, Fieldclass Ticket(BaseModel):    category: Literal["billing", "technical", "account"]    priority: Literal["low", "medium", "high"]    summary: str = Field(max_length=200, description="One sentence, no customer names")    customer_id: str | None = Field(default=None, description="null if not mentioned")chain = prompt | llm.with_structured_output(Ticket, method="json_schema", strict=True)ticket = chain.invoke({"message": text})          # a Ticket instance, not a string

On ChatOpenAI, method="json_schema" uses the provider's structured-output mode; other chat model classes pick the best mechanism their provider supports, often tool calling.

Three details that decide whether it holds

DetailWhy it matters
Literal for categoriesA free str lets the model drift to "Billing issue" or "payments"
Optional fields (None allowed) where data may be absentA required field with no real value forces the model to invent one
Field(description=...)Descriptions are sent to the model; put disambiguation rules here, next to the field

Keep a bounded repair path

  1. Constrain — with_structured_output on providers that support it.
  2. Retry once on transient errors — .with_retry(stop_after_attempt=2) for network or rate-limit failures.
  3. Inspect failures — include_raw=True returns the raw message and any parsing error alongside the parsed object, so you can log it.
  4. Fallback for weaker models — PydanticOutputParser with one OutputFixingParser attempt (from langchain-classic), then a review queue.

More than one repair attempt is usually wasted spend: if the model fails twice on the same input, the input is unusual and a person should see it.

In an agent

With LangChain 1.x create_agent, pass response_format=Ticket and the agent's final answer is returned as structured data in the result's structured_response, after any tool calls.

The residual risk

Schema-valid but wrong output, such as priority: "low" for an outage report, passes every parser. So keep two metrics apart: parse-failure rate (should be near zero) and field accuracy on a labelled LangSmith dataset, which is the real quality measure.

A real-life example

Scenario (illustrative numbers). A SaaS helpdesk classifies 50,000 incoming emails a day into category and priority using a prompt plus JsonOutputParser. About 1.5% fail to parse, and the downstream router drops them silently, so some urgent tickets wait a day.

The team moves to with_structured_output(Ticket) with Literal fields and include_raw=True for logging. Parse failures fall to about 0.01%, all logged and sent to a manual queue. The LangSmith dataset of 400 labelled emails then shows priority accuracy at only 83%: "high" was being given to angry but minor tickets. Adding a priority rule to the field description lifts it to 91%, a problem the parse errors had been hiding.

Follow-up questions to expect

  • "What's the difference between json_mode and json_schema?" — JSON mode guarantees valid JSON of any shape; JSON-schema mode constrains the output to your specific schema.
  • "Can you stream structured output?" — Yes, partial objects can be streamed on supported models, but validate only the final object.
  • "How do you handle nested or list outputs?" — Nested Pydantic models and list[Item] fields work; keep nesting shallow, since very complex schemas may be rejected by strict modes.