LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you validate inputs for LangChain chains?


What you need to know

Why validate before the model

  • Cost — a 200-page paste becomes a very expensive call, or a context-length error after a long wait.
  • Correctness — a missing employee_id gives a fluent answer about the wrong person.
  • Security — untrusted text can try to override your instructions.
  • Clear errors — a Pydantic error says exactly which field is wrong; a model failure says nothing.

The pattern

Python
from pydantic import BaseModel, Field, field_validatorfrom langchain_core.runnables import RunnableLambdaclass LeaveQuestion(BaseModel):    question: str = Field(min_length=3, max_length=2000)    employee_id: str = Field(pattern=r"^E\d{5}$")    top_k: int = Field(default=4, ge=1, le=10)    @field_validator("question")    @classmethod    def strip_text(cls, v: str) -> str:        return v.strip()def validate(raw: dict) -> dict:    return LeaveQuestion.model_validate(raw).model_dump()safe_chain = RunnableLambda(validate).with_config(run_name="validate_input") | chainsafe_chain.invoke({"question": "How many leaves are left?", "employee_id": "E10233"})
  • model_validate raises pydantic.ValidationError listing every failing field.
  • Because the validator is a runnable, it shows as a named step (validate_input) in the trace.
  • The API layer (for example FastAPI) catches ValidationError and returns a 422 with the field errors.

You can also declare the chain's input type with chain.with_types(input_type=LeaveQuestion). That changes the chain's published schema (get_input_jsonschema()), which is useful for documentation and serving, but it does not by itself reject bad input at runtime — the explicit validation step does.

Validating tool inputs in agents

When you define a tool with @tool and type hints, or with an args_schema Pydantic model, the model's tool-call arguments are validated before the function runs. If validation fails, the agent's tool node returns an error message to the model so it can correct itself. Add business rules inside the tool too — "leave days must be between 0.5 and 30".

Safety checks for untrusted text

  • Cap length (characters or tokens).
  • Put user text in a clearly delimited place in the prompt, never inside the system message.
  • Do not interpolate user text into tool names, SQL or file paths.
  • For agents, consider PIIMiddleware to redact emails or card numbers before they reach the model.

A real-life example

An HR bot files leave requests. Its apply_leave tool first accepted days: str. The model once sent "days": "two and a half", and the HR system stored 0. Another time a user asked for "leave from 3rd to 1st" and the system created a negative-length leave.

The team changed the tool to a Pydantic schema: start: date, end: date, days: float = Field(gt=0, le=30), with a validator that end >= start. Invalid calls now come back to the model as a clear error ("end must be on or after start"), the model asks the employee to confirm dates, and bad records in the HR system dropped to zero.

Follow-up questions to expect

  • "Is Pydantic v1 syntax still fine?" — No; use v2: model_validate, model_dump, field_validator. LangChain 1.x uses Pydantic 2.
  • "Can validation stop prompt injection?" — It reduces risk (length, format, delimiting) but cannot detect every injection; limit what tools can do and require confirmation for risky actions.
  • "Where should validation live — API or chain?" — Both is fine; the chain-level step protects every caller, including batch jobs.