LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Implement a function to validate LangChain output formats.


What you need to know

Three levels of checking

  1. Syntax — is it valid JSON at all?
  2. Schema — right fields, types, allowed values (Pydantic does this).
  3. Semantics — do the values make sense for the business? (Your code does this.)

with_structured_output handles 1 and 2. Only your code can do 3.

The function

Python
import loggingfrom typing import Literalfrom pydantic import BaseModel, Fieldlog = logging.getLogger(__name__)class LeaveDecision(BaseModel):    action: Literal["approve", "reject", "escalate"]    days: float = Field(gt=0, le=30)    reason: str = Field(max_length=300)structured = llm.with_structured_output(LeaveDecision, include_raw=True)def decide(request_text: str, balance: float) -> LeaveDecision | None:    out = structured.invoke(request_text)          # {"raw", "parsed", "parsing_error"}    decision = out["parsed"]    if out["parsing_error"] or decision is None:        log.warning("format_failure", extra={"raw": str(out["raw"].content)[:300]})        return None                                  # caller retries once or escalates    if decision.action == "approve" and decision.days > balance:        return decision.model_copy(update={"action": "escalate",                                           "reason": "Not enough leave balance"})    return decision
  • Literal[...] restricts values; Field(gt=0, le=30) checks ranges.
  • include_raw=True returns the raw message and any parsing error instead of raising, so you can log and decide.
  • The balance check is semantic validation — the schema cannot know the employee's balance.

with_structured_output accepts a method argument on many providers ("json_schema", "function_calling", "json_mode"); the provider-native JSON-schema mode is the most reliable where supported.

When the provider has no structured output

Use a parser with format instructions in the prompt: PydanticOutputParser from langchain_core.output_parsers. For repair, OutputFixingParser and RetryWithErrorOutputParser send the bad output back to a model; in LangChain 1.x they are in langchain_classic.output_parsers.

In agents

Python
from langchain.agents import create_agentfrom langchain.agents.structured_output import ToolStrategyagent = create_agent(model, tools=[...], response_format=ToolStrategy(LeaveDecision))result = agent.invoke({"messages": [...]})result["structured_response"]      # a validated LeaveDecision

ToolStrategy asks for the answer through a tool call and, by default, sends validation errors back to the model to fix. ProviderStrategy uses the provider's native structured output.

A real-life example

An HR bot reads leave emails and proposes a decision for a manager. Version one asked for "JSON with action, days and reason" and used json.loads. About 3% of replies had text before the JSON, and once the model returned "days": "3 days", which broke the payroll import.

Version two uses with_structured_output(LeaveDecision). Format failures dropped to under 0.1%. A semantic check then caught a new class of bug the schema could not: the model approving 12 days for someone with 4 left. Those now become escalate with a reason, and the manager sees them flagged.

Follow-up questions to expect

  • "Structured output vs output parser?" — Structured output uses the provider's constrained mode or tool calling; a parser trusts the model to follow prompt instructions and fixes it afterwards. Prefer structured output.
  • "How do you handle optional fields?" — Make them Optional with defaults, and describe when they should be empty; forced fields make the model invent values.
  • "Does validation guarantee correctness?" — No. It guarantees shape; correctness needs business checks and evaluation.