Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 4: Structured Output Instability
Scenario: a LangChain chain feeds JSON into a downstream service, and occasionally the model returns prose or a fenced block and parsing fails. How do you make the output reliable?
What you need to know
An output parser on raw text is a repair step: the model writes whatever it likes, and you try to fix it. with_structured_output is a prevention step: the model is asked, through the provider's native mechanism, to produce exactly your schema.
The modern pattern
1from typing import Literal2from pydantic import BaseModel, Field34class Ticket(BaseModel):5 category: Literal["billing", "technical", "account"]6 priority: Literal["low", "medium", "high"]7 summary: str = Field(max_length=200, description="One sentence, no customer names")8 customer_id: str | None = Field(default=None, description="null if not mentioned")910chain = prompt | llm.with_structured_output(Ticket, method="json_schema", strict=True)11ticket = chain.invoke({"message": text}) # a Ticket instance, not a stringOn ChatOpenAI, method="json_schema" uses the provider's structured-output mode; other chat model classes pick the best mechanism their provider supports, often tool calling.
Three details that decide whether it holds
| Detail | Why it matters |
|---|---|
Literal for categories | A free str lets the model drift to "Billing issue" or "payments" |
Optional fields (None allowed) where data may be absent | A required field with no real value forces the model to invent one |
Field(description=...) | Descriptions are sent to the model; put disambiguation rules here, next to the field |
Keep a bounded repair path
- Constrain —
with_structured_outputon providers that support it. - Retry once on transient errors —
.with_retry(stop_after_attempt=2)for network or rate-limit failures. - Inspect failures —
include_raw=Truereturns the raw message and any parsing error alongside the parsed object, so you can log it. - Fallback for weaker models —
PydanticOutputParserwith oneOutputFixingParserattempt (fromlangchain-classic), then a review queue.
More than one repair attempt is usually wasted spend: if the model fails twice on the same input, the input is unusual and a person should see it.
In an agent
With LangChain 1.x create_agent, pass response_format=Ticket and the agent's final answer is returned as structured data in the result's structured_response, after any tool calls.
The residual risk
Schema-valid but wrong output, such as priority: "low" for an outage report, passes every parser. So keep two metrics apart: parse-failure rate (should be near zero) and field accuracy on a labelled LangSmith dataset, which is the real quality measure.
A real-life example
Scenario (illustrative numbers). A SaaS helpdesk classifies 50,000 incoming emails a day into category and priority using a prompt plus JsonOutputParser. About 1.5% fail to parse, and the downstream router drops them silently, so some urgent tickets wait a day.
The team moves to with_structured_output(Ticket) with Literal fields and include_raw=True for logging. Parse failures fall to about 0.01%, all logged and sent to a manual queue. The LangSmith dataset of 400 labelled emails then shows priority accuracy at only 83%: "high" was being given to angry but minor tickets. Adding a priority rule to the field description lifts it to 91%, a problem the parse errors had been hiding.
Follow-up questions to expect
- "What's the difference between json_mode and json_schema?" — JSON mode guarantees valid JSON of any shape; JSON-schema mode constrains the output to your specific schema.
- "Can you stream structured output?" — Yes, partial objects can be streamed on supported models, but validate only the final object.
- "How do you handle nested or list outputs?" — Nested Pydantic models and
list[Item]fields work; keep nesting shallow, since very complex schemas may be rejected by strict modes.