LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to chain text generation and parsing.


What you need to know

"Generation then parsing" is the most common chain in production: the model writes, the parser turns it into data your code can store or act on.

Python
from typing import Literalfrom pydantic import BaseModel, Fieldfrom langchain_core.prompts import ChatPromptTemplateclass LeadInfo(BaseModel):    """Details extracted from a property enquiry."""    budget_inr_lakh: float | None = Field(description="Budget in lakh rupees, if stated")    locality: str | None    bhk: int | None = Field(ge=1, le=6)    move_in_months: int | None = Field(description="Months until they want to move")    intent: Literal["buy", "rent", "unclear"]def build_lead_extractor(llm):    prompt = ChatPromptTemplate.from_messages([        ("system", "Extract the fields from the enquiry. Use null when a "                   "field is not stated. Do not guess."),        ("human", "{enquiry}"),    ])    return (prompt | llm.with_structured_output(LeadInfo)).with_retry(stop_after_attempt=2)extract = build_lead_extractor(llm)info = extract.invoke({"enquiry": "Need a 2BHK in Whitefield, around 85 lakh, by March"})info.bhk, info.budget_inr_lakh        # 2, 85.0
  • Field descriptions and the class docstring become part of the schema the model sees, so write them for the model.
  • | None with "use null when not stated" stops the model inventing a budget.
  • Literal restricts intent to three values.

Without structured output support

Python
from langchain_core.output_parsers import PydanticOutputParserparser = PydanticOutputParser(pydantic_object=LeadInfo)prompt = prompt.partial(format_instructions=parser.get_format_instructions())chain = prompt | llm | parser         # the system message must include {format_instructions}

This asks for JSON in words and validates after. It works on any model but fails more often.

Failure handling

  • Parse or validation error: retry once, ideally with the error message added to the prompt; after that, log and send to review.
  • Values that pass the schema but are wrong: add business checks (a budget of 0.85 lakh for a flat is probably a unit mistake).

A real-life example

A real-estate firm's lead-qualification pipeline extracted fields with a prompt that said "return JSON" and a json.loads call. Around 1 in 40 enquiries broke the job: extra text around the JSON, a budget written as "85L", or a missing field. After switching to with_structured_output(LeadInfo), malformed output almost disappeared. The remaining issue was units: "1.2 Cr" came back as 1.2 in the lakh field. The team fixed it in the field description ("convert crore to lakh: 1 crore = 100 lakh") and added a check that flags any flat budget under 10 lakh for review. On 500 labelled enquiries, field accuracy rose from 88% to 96%.

Follow-up questions to expect

  • "Why describe the fields?" — The descriptions are sent with the schema, and they are the main way to tell the model what each field means.
  • "How do you get the raw reply too?" — Use with_structured_output(LeadInfo, include_raw=True), which returns the raw message, the parsed object and any parsing error.
  • "TypedDict or Pydantic?" — Pydantic validates at run time; a TypedDict or plain JSON schema returns a dict with no validation.