Course Content
LangChain Mastery
7 sections · 109 lessons
How do you validate inputs for LangChain chains?
What you need to know
Why validate before the model
- Cost — a 200-page paste becomes a very expensive call, or a context-length error after a long wait.
- Correctness — a missing
employee_idgives a fluent answer about the wrong person. - Security — untrusted text can try to override your instructions.
- Clear errors — a Pydantic error says exactly which field is wrong; a model failure says nothing.
The pattern
1from pydantic import BaseModel, Field, field_validator2from langchain_core.runnables import RunnableLambda34class LeaveQuestion(BaseModel):5 question: str = Field(min_length=3, max_length=2000)6 employee_id: str = Field(pattern=r"^E\d{5}$")7 top_k: int = Field(default=4, ge=1, le=10)89 @field_validator("question")10 @classmethod11 def strip_text(cls, v: str) -> str:12 return v.strip()1314def validate(raw: dict) -> dict:15 return LeaveQuestion.model_validate(raw).model_dump()1617safe_chain = RunnableLambda(validate).with_config(run_name="validate_input") | chain18safe_chain.invoke({"question": "How many leaves are left?", "employee_id": "E10233"})model_validateraisespydantic.ValidationErrorlisting every failing field.- Because the validator is a runnable, it shows as a named step (
validate_input) in the trace. - The API layer (for example FastAPI) catches
ValidationErrorand returns a 422 with the field errors.
You can also declare the chain's input type with chain.with_types(input_type=LeaveQuestion). That changes the chain's published schema (get_input_jsonschema()), which is useful for documentation and serving, but it does not by itself reject bad input at runtime — the explicit validation step does.
Validating tool inputs in agents
When you define a tool with @tool and type hints, or with an args_schema Pydantic model, the model's tool-call arguments are validated before the function runs. If validation fails, the agent's tool node returns an error message to the model so it can correct itself. Add business rules inside the tool too — "leave days must be between 0.5 and 30".
Safety checks for untrusted text
- Cap length (characters or tokens).
- Put user text in a clearly delimited place in the prompt, never inside the system message.
- Do not interpolate user text into tool names, SQL or file paths.
- For agents, consider
PIIMiddlewareto redact emails or card numbers before they reach the model.
A real-life example
An HR bot files leave requests. Its apply_leave tool first accepted days: str. The model once sent "days": "two and a half", and the HR system stored 0. Another time a user asked for "leave from 3rd to 1st" and the system created a negative-length leave.
The team changed the tool to a Pydantic schema: start: date, end: date, days: float = Field(gt=0, le=30), with a validator that end >= start. Invalid calls now come back to the model as a clear error ("end must be on or after start"), the model asks the employee to confirm dates, and bad records in the HR system dropped to zero.
Follow-up questions to expect
- "Is Pydantic v1 syntax still fine?" — No; use v2:
model_validate,model_dump,field_validator. LangChain 1.x uses Pydantic 2. - "Can validation stop prompt injection?" — It reduces risk (length, format, delimiting) but cannot detect every injection; limit what tools can do and require confirmation for risky actions.
- "Where should validation live — API or chain?" — Both is fine; the chain-level step protects every caller, including batch jobs.