Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Validating with Pydantic
The provider now guarantees ShipFast gets JSON in the right shape. That is not the same as data you can act on. The rescheduling system will book a slot for whatever date you give it. If the model says 2025-01-01, or copies a tracking ID the customer never wrote, the booking goes wrong and a customer waits for a parcel that is not coming.
So every model reply passes through a second gate: a Pydantic model that checks the values, not just the shape. Think of the model's reply the way you think of a request body from the public internet. You would never write it to the database without validation. The same rule applies here, with one extra twist: some checks need to see the customer's original message.
Two models, one file
ShipFast's classification and extraction outputs live in shipfast/schemas.py. Here is the classification part.
1# shipfast/schemas.py (part 1)2import re3from datetime import date, timedelta4from enum import StrEnum5from typing import Literal67from pydantic import BaseModel, ConfigDict, Field, ValidationInfo, field_validator8910class Intent(StrEnum):11 RESCHEDULE = "reschedule"12 ADDRESS_CHANGE = "address_change"13 DAMAGED_PARCEL = "damaged_parcel"14 COMPLAINT = "complaint"15 OTHER = "other"16 UNKNOWN = "unknown"171819class Classification(BaseModel):20 model_config = ConfigDict(extra="forbid")21 reason: str = Field(description="One short sentence, at most 20 words")22 intent: Intent2324 @field_validator("intent", mode="before")25 @classmethod26 def normalise_label(cls, value: object) -> object:27 if isinstance(value, str):28 return value.strip().lower().replace(" ", "_").replace("-", "_")29 return valueextra="forbid" rejects keys you did not ask for and makes the generated schema say additionalProperties: false, which providers need for strict mode. The mode="before" validator runs before type checking, so "Address Change" and "address-change" become address_change. It is a cheap, deterministic fix for a harmless difference. Anything that still does not match the enum, such as "refund", is rejected.
Required, but allowed to be null
Extraction has six fields, and most messages mention only two or three. The tempting design is optional fields with defaults. The better one is required fields that may be null.
1# shipfast/schemas.py (part 2)2HHMM = re.compile(r"([01]\d|2[0-3]):[0-5]\d")3TRACKING = re.compile(r"SF\d{8}")456class Extraction(BaseModel):7 model_config = ConfigDict(extra="forbid")8 tracking_id: str | None = Field(description="Like SF12345678, copied exactly, or null")9 delivery_date: date | None = Field(description="ISO date the customer asks for, or null")10 window_start: str | None = Field(description="24-hour HH:MM, or null")11 window_end: str | None = Field(description="24-hour HH:MM, or null")12 new_address: str | None = Field(description="Full new address as written, or null")13 pincode: str | None = Field(description="6-digit PIN code of the new address, or null")None of the fields has a default, so all six are in the schema's required list. The model must write every key and choose, for each one, a value or an explicit null. With optional fields, a missing key could mean "the customer did not say" or "the model forgot to look", and you cannot tell which. With required nullable fields, null is a decision the model made. Strict schema enforcement also expects every property to be required, so this design works with the providers rather than against them.
The descriptions matter too. They are part of the schema the model reads, so "copied exactly" and "or null" are instructions placed right next to the field they apply to.
Validators that check the facts
Now the checks that catch the failures structured outputs cannot. They go inside the Extraction class.
1 # shipfast/schemas.py, inside class Extraction2 @field_validator("tracking_id")3 @classmethod4 def tracking_id_is_real(cls, v: str | None, info: ValidationInfo) -> str | None:5 if v is None:6 return None7 v = v.upper().replace(" ", "")8 if not TRACKING.fullmatch(v):9 raise ValueError("must look like SF12345678")10 source = (info.context or {}).get("source", "").upper().replace(" ", "")11 if v not in source:12 raise ValueError("does not appear in the customer's message")13 return v1415 @field_validator("delivery_date")16 @classmethod17 def date_in_range(cls, v: date | None, info: ValidationInfo) -> date | None:18 today = (info.context or {}).get("today")19 if v is not None and today is not None and not today <= v <= today + timedelta(days=14):20 raise ValueError(f"must be between {today} and 14 days later")21 return v2223 @field_validator("window_start", "window_end")24 @classmethod25 def is_hhmm(cls, v: str | None) -> str | None:26 if v is not None and not HHMM.fullmatch(v):27 raise ValueError("must be HH:MM in 24-hour time")28 return v2930 @field_validator("pincode")31 @classmethod32 def six_digits(cls, v: str | None) -> str | None:33 if v is not None and not re.fullmatch(r"[1-9]\d{5}", v):34 raise ValueError("must be a 6-digit PIN code")35 return vThe most valuable check is the first one. A model asked to extract a tracking ID will sometimes produce a perfectly formatted ID that is not in the message, because it has seen thousands of IDs that look like that. This is called a grounding check: every value that should be copied from the input must actually appear in the input. It costs one substring search and catches a failure that no schema can.
Passing context into validation
The grounding check needs the customer's message, and the date check needs today's date. Neither belongs in the model's output. Pydantic v2 lets you pass them at validation time with context.
1from datetime import date23from pydantic import ValidationError45from shipfast.schemas import Extraction67reply = ('{"tracking_id": "SF99999999", "delivery_date": "2025-01-01", "window_start": "6pm",'8 ' "window_end": null, "new_address": null, "pincode": null}')9try:10 Extraction.model_validate_json(11 reply, context={"source": "SF12345678 deliver tomorrow", "today": date(2026, 9, 23)})12except ValidationError as err:13 for e in err.errors():14 print(e["loc"], e["msg"])15# ('tracking_id',) Value error, does not appear in the customer's message16# ('delivery_date',) Value error, must be between 2026-09-23 and 14 days later17# ('window_start',) Value error, must be HH:MM in 24-hour timePydantic collects every error, not just the first. That list is useful twice: in your logs, where it tells you which field fails most often, and in the next lesson, where it becomes the message you send back to the model to ask for a correction.
Lax parsing, on purpose
Pydantic v2 has two modes. In the default lax mode, it converts compatible values: the string "2026-09-24" becomes a date, and the string "5" would become an integer 5. In strict mode it refuses any conversion.
For model replies, lax mode is what you want for dates, because JSON has no date type and the model can only send a string. But know what lax mode accepts. It would also accept "true" for a boolean field and 1.0 for an integer. For ShipFast's schema this is harmless. If you add a field where a conversion could hide a real mistake, such as an amount in rupees, mark that one field strict with Field(strict=True) and keep the rest lax.
The general rule: be lax about how a value is written and strict about what it means. "Address Change" and address_change mean the same thing, so normalise. A date in the past means something wrong, so reject.
Check your understanding
0 of 3 answered
1.Why does ShipFast make extraction fields required but nullable, instead of optional with a default of None?
2.The model returns tracking_id: "SF48213377", correctly formatted, but the customer's message contains SF48213372. Which check catches this?
3.Where should today's date come from when validating delivery_date?