Course Content
LangChain Mastery
7 sections · 109 lessons
Implement a function to validate LangChain output formats.
What you need to know
Three levels of checking
- Syntax — is it valid JSON at all?
- Schema — right fields, types, allowed values (Pydantic does this).
- Semantics — do the values make sense for the business? (Your code does this.)
with_structured_output handles 1 and 2. Only your code can do 3.
The function
1import logging2from typing import Literal3from pydantic import BaseModel, Field45log = logging.getLogger(__name__)67class LeaveDecision(BaseModel):8 action: Literal["approve", "reject", "escalate"]9 days: float = Field(gt=0, le=30)10 reason: str = Field(max_length=300)1112structured = llm.with_structured_output(LeaveDecision, include_raw=True)1314def decide(request_text: str, balance: float) -> LeaveDecision | None:15 out = structured.invoke(request_text) # {"raw", "parsed", "parsing_error"}16 decision = out["parsed"]17 if out["parsing_error"] or decision is None:18 log.warning("format_failure", extra={"raw": str(out["raw"].content)[:300]})19 return None # caller retries once or escalates20 if decision.action == "approve" and decision.days > balance:21 return decision.model_copy(update={"action": "escalate",22 "reason": "Not enough leave balance"})23 return decisionLiteral[...]restricts values;Field(gt=0, le=30)checks ranges.include_raw=Truereturns the raw message and any parsing error instead of raising, so you can log and decide.- The balance check is semantic validation — the schema cannot know the employee's balance.
with_structured_output accepts a method argument on many providers ("json_schema", "function_calling", "json_mode"); the provider-native JSON-schema mode is the most reliable where supported.
When the provider has no structured output
Use a parser with format instructions in the prompt: PydanticOutputParser from langchain_core.output_parsers. For repair, OutputFixingParser and RetryWithErrorOutputParser send the bad output back to a model; in LangChain 1.x they are in langchain_classic.output_parsers.
In agents
1from langchain.agents import create_agent2from langchain.agents.structured_output import ToolStrategy34agent = create_agent(model, tools=[...], response_format=ToolStrategy(LeaveDecision))5result = agent.invoke({"messages": [...]})6result["structured_response"] # a validated LeaveDecisionToolStrategy asks for the answer through a tool call and, by default, sends validation errors back to the model to fix. ProviderStrategy uses the provider's native structured output.
A real-life example
An HR bot reads leave emails and proposes a decision for a manager. Version one asked for "JSON with action, days and reason" and used json.loads. About 3% of replies had text before the JSON, and once the model returned "days": "3 days", which broke the payroll import.
Version two uses with_structured_output(LeaveDecision). Format failures dropped to under 0.1%. A semantic check then caught a new class of bug the schema could not: the model approving 12 days for someone with 4 left. Those now become escalate with a reason, and the manager sees them flagged.
Follow-up questions to expect
- "Structured output vs output parser?" — Structured output uses the provider's constrained mode or tool calling; a parser trusts the model to follow prompt instructions and fixes it afterwards. Prefer structured output.
- "How do you handle optional fields?" — Make them
Optionalwith defaults, and describe when they should be empty; forced fields make the model invent values. - "Does validation guarantee correctness?" — No. It guarantees shape; correctness needs business checks and evaluation.