Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to chain text generation and parsing.
What you need to know
"Generation then parsing" is the most common chain in production: the model writes, the parser turns it into data your code can store or act on.
1from typing import Literal2from pydantic import BaseModel, Field3from langchain_core.prompts import ChatPromptTemplate45class LeadInfo(BaseModel):6 """Details extracted from a property enquiry."""7 budget_inr_lakh: float | None = Field(description="Budget in lakh rupees, if stated")8 locality: str | None9 bhk: int | None = Field(ge=1, le=6)10 move_in_months: int | None = Field(description="Months until they want to move")11 intent: Literal["buy", "rent", "unclear"]1213def build_lead_extractor(llm):14 prompt = ChatPromptTemplate.from_messages([15 ("system", "Extract the fields from the enquiry. Use null when a "16 "field is not stated. Do not guess."),17 ("human", "{enquiry}"),18 ])19 return (prompt | llm.with_structured_output(LeadInfo)).with_retry(stop_after_attempt=2)2021extract = build_lead_extractor(llm)22info = extract.invoke({"enquiry": "Need a 2BHK in Whitefield, around 85 lakh, by March"})23info.bhk, info.budget_inr_lakh # 2, 85.0- Field descriptions and the class docstring become part of the schema the model sees, so write them for the model.
| Nonewith "use null when not stated" stops the model inventing a budget.Literalrestrictsintentto three values.
Without structured output support
1from langchain_core.output_parsers import PydanticOutputParser23parser = PydanticOutputParser(pydantic_object=LeadInfo)4prompt = prompt.partial(format_instructions=parser.get_format_instructions())5chain = prompt | llm | parser # the system message must include {format_instructions}This asks for JSON in words and validates after. It works on any model but fails more often.
Failure handling
- Parse or validation error: retry once, ideally with the error message added to the prompt; after that, log and send to review.
- Values that pass the schema but are wrong: add business checks (a budget of 0.85 lakh for a flat is probably a unit mistake).
A real-life example
A real-estate firm's lead-qualification pipeline extracted fields with a prompt that said "return JSON" and a json.loads call. Around 1 in 40 enquiries broke the job: extra text around the JSON, a budget written as "85L", or a missing field. After switching to with_structured_output(LeadInfo), malformed output almost disappeared. The remaining issue was units: "1.2 Cr" came back as 1.2 in the lakh field. The team fixed it in the field description ("convert crore to lakh: 1 crore = 100 lakh") and added a check that flags any flat budget under 10 lakh for review. On 500 labelled enquiries, field accuracy rose from 88% to 96%.
Follow-up questions to expect
- "Why describe the fields?" — The descriptions are sent with the schema, and they are the main way to tell the model what each field means.
- "How do you get the raw reply too?" — Use
with_structured_output(LeadInfo, include_raw=True), which returns the raw message, the parsed object and any parsing error. - "TypedDict or Pydantic?" — Pydantic validates at run time; a
TypedDictor plain JSON schema returns a dict with no validation.