Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to validate LangChain LLM outputs.
What you need to know
An LLM can return valid-looking JSON with wrong values, or text that is not JSON at all. Validation happens at three levels.
| Level | Question it answers | Tool |
|---|---|---|
| Shape | Is it the right structure and types? | Pydantic model, structured output |
| Rules | Do the values make sense together? | Pydantic validators, your own checks |
| Grounding | Are the values actually in the source? | String or number match against the input |
1from datetime import date2from pydantic import BaseModel, Field, model_validator34class Invoice(BaseModel):5 vendor: str6 gstin: str = Field(pattern=r"^\d{2}[A-Z]{5}\d{4}[A-Z][1-9A-Z]Z[0-9A-Z]$")7 invoice_number: str8 invoice_date: date9 subtotal: float = Field(ge=0)10 gst: float = Field(ge=0)11 total: float = Field(gt=0)1213 @model_validator(mode="after")14 def totals_add_up(self):15 if abs(self.subtotal + self.gst - self.total) > 1:16 raise ValueError("subtotal + gst does not equal total")17 return self1819extractor = prompt | llm.with_structured_output(Invoice, include_raw=True)2021def extract_invoice(text: str) -> Invoice | None:22 result = extractor.invoke({"text": text})23 if result["parsing_error"] is not None:24 return None # send to review queue25 inv = result["parsed"]26 if inv.invoice_number not in text: # grounding check27 return None28 return invThe Field constraints check each value (GSTIN format, no negative amounts). The model_validator checks values together. With include_raw=True a failure does not raise; you get parsing_error and the raw message, which you can log. The last check makes sure the invoice number was copied from the document, not invented.
When structured output is not available
Use PydanticOutputParser: put parser.get_format_instructions() in the prompt, then prompt | llm | parser. It raises OutputParserException on bad output. OutputFixingParser, which sends bad output back to the model for repair, now lives in langchain_classic.output_parsers; a simple with_retry or one manual retry with the error message usually does the same job.
A real-life example
An invoice extraction chain at a Chennai logistics firm handles 900 invoices a week. Before validation, about 1 in 30 records had a total that did not match subtotal plus GST, often because the model read a line-item amount as the total. After adding the Invoice model with the totals rule and the GSTIN pattern, those records were caught before reaching the accounting system. About 3% of invoices go to a human queue, and each one arrives with the raw model output and the exact rule that failed, so a clerk fixes it in under a minute.
Follow-up questions to expect
- "
with_structured_outputor an output parser?" — Structured output when the provider supports it, because generation itself is constrained; a parser for models that don't. - "What does
include_raw=Truegive you?" — A dict withraw(theAIMessage),parsed(the object orNone) andparsing_error, so failures can be logged instead of crashing. - "Should you retry on validation failure?" — Once, with the error message in the prompt. If it fails twice, send it to a person; more retries rarely fix a real misreading.