Building AI Features in Python Backends

Validating with Pydantic


The provider now guarantees ShipFast gets JSON in the right shape. That is not the same as data you can act on. The rescheduling system will book a slot for whatever date you give it. If the model says 2025-01-01, or copies a tracking ID the customer never wrote, the booking goes wrong and a customer waits for a parcel that is not coming.

So every model reply passes through a second gate: a Pydantic model that checks the values, not just the shape. Think of the model's reply the way you think of a request body from the public internet. You would never write it to the database without validation. The same rule applies here, with one extra twist: some checks need to see the customer's original message.

From schema-shaped reply to data you can bookReply matchesthe JSON schemaTypes and enumvalues checkedTracking IDfound inthe message?Date within 14days of today?ValidatedExtractionobjectValidation context carries the message and today's date.
The provider guarantees shape; only checks against the original message catch an invented ID or last Monday's date.

Two models, one file

ShipFast's classification and extraction outputs live in shipfast/schemas.py. Here is the classification part.

Python
# shipfast/schemas.py (part 1)import refrom datetime import date, timedeltafrom enum import StrEnumfrom typing import Literalfrom pydantic import BaseModel, ConfigDict, Field, ValidationInfo, field_validatorclass Intent(StrEnum):    RESCHEDULE = "reschedule"    ADDRESS_CHANGE = "address_change"    DAMAGED_PARCEL = "damaged_parcel"    COMPLAINT = "complaint"    OTHER = "other"    UNKNOWN = "unknown"class Classification(BaseModel):    model_config = ConfigDict(extra="forbid")    reason: str = Field(description="One short sentence, at most 20 words")    intent: Intent    @field_validator("intent", mode="before")    @classmethod    def normalise_label(cls, value: object) -> object:        if isinstance(value, str):            return value.strip().lower().replace(" ", "_").replace("-", "_")        return value

extra="forbid" rejects keys you did not ask for and makes the generated schema say additionalProperties: false, which providers need for strict mode. The mode="before" validator runs before type checking, so "Address Change" and "address-change" become address_change. It is a cheap, deterministic fix for a harmless difference. Anything that still does not match the enum, such as "refund", is rejected.

Required, but allowed to be null

Extraction has six fields, and most messages mention only two or three. The tempting design is optional fields with defaults. The better one is required fields that may be null.

Python
# shipfast/schemas.py (part 2)HHMM = re.compile(r"([01]\d|2[0-3]):[0-5]\d")TRACKING = re.compile(r"SF\d{8}")class Extraction(BaseModel):    model_config = ConfigDict(extra="forbid")    tracking_id: str | None = Field(description="Like SF12345678, copied exactly, or null")    delivery_date: date | None = Field(description="ISO date the customer asks for, or null")    window_start: str | None = Field(description="24-hour HH:MM, or null")    window_end: str | None = Field(description="24-hour HH:MM, or null")    new_address: str | None = Field(description="Full new address as written, or null")    pincode: str | None = Field(description="6-digit PIN code of the new address, or null")

None of the fields has a default, so all six are in the schema's required list. The model must write every key and choose, for each one, a value or an explicit null. With optional fields, a missing key could mean "the customer did not say" or "the model forgot to look", and you cannot tell which. With required nullable fields, null is a decision the model made. Strict schema enforcement also expects every property to be required, so this design works with the providers rather than against them.

The descriptions matter too. They are part of the schema the model reads, so "copied exactly" and "or null" are instructions placed right next to the field they apply to.

Validators that check the facts

Now the checks that catch the failures structured outputs cannot. They go inside the Extraction class.

Python
    # shipfast/schemas.py, inside class Extraction    @field_validator("tracking_id")    @classmethod    def tracking_id_is_real(cls, v: str | None, info: ValidationInfo) -> str | None:        if v is None:            return None        v = v.upper().replace(" ", "")        if not TRACKING.fullmatch(v):            raise ValueError("must look like SF12345678")        source = (info.context or {}).get("source", "").upper().replace(" ", "")        if v not in source:            raise ValueError("does not appear in the customer's message")        return v    @field_validator("delivery_date")    @classmethod    def date_in_range(cls, v: date | None, info: ValidationInfo) -> date | None:        today = (info.context or {}).get("today")        if v is not None and today is not None and not today <= v <= today + timedelta(days=14):            raise ValueError(f"must be between {today} and 14 days later")        return v    @field_validator("window_start", "window_end")    @classmethod    def is_hhmm(cls, v: str | None) -> str | None:        if v is not None and not HHMM.fullmatch(v):            raise ValueError("must be HH:MM in 24-hour time")        return v    @field_validator("pincode")    @classmethod    def six_digits(cls, v: str | None) -> str | None:        if v is not None and not re.fullmatch(r"[1-9]\d{5}", v):            raise ValueError("must be a 6-digit PIN code")        return v

The most valuable check is the first one. A model asked to extract a tracking ID will sometimes produce a perfectly formatted ID that is not in the message, because it has seen thousands of IDs that look like that. This is called a grounding check: every value that should be copied from the input must actually appear in the input. It costs one substring search and catches a failure that no schema can.

Passing context into validation

The grounding check needs the customer's message, and the date check needs today's date. Neither belongs in the model's output. Pydantic v2 lets you pass them at validation time with context.

Python
from datetime import datefrom pydantic import ValidationErrorfrom shipfast.schemas import Extractionreply = ('{"tracking_id": "SF99999999", "delivery_date": "2025-01-01", "window_start": "6pm",'         ' "window_end": null, "new_address": null, "pincode": null}')try:    Extraction.model_validate_json(        reply, context={"source": "SF12345678 deliver tomorrow", "today": date(2026, 9, 23)})except ValidationError as err:    for e in err.errors():        print(e["loc"], e["msg"])# ('tracking_id',) Value error, does not appear in the customer's message# ('delivery_date',) Value error, must be between 2026-09-23 and 14 days later# ('window_start',) Value error, must be HH:MM in 24-hour time

Pydantic collects every error, not just the first. That list is useful twice: in your logs, where it tells you which field fails most often, and in the next lesson, where it becomes the message you send back to the model to ask for a correction.

Lax parsing, on purpose

Pydantic v2 has two modes. In the default lax mode, it converts compatible values: the string "2026-09-24" becomes a date, and the string "5" would become an integer 5. In strict mode it refuses any conversion.

For model replies, lax mode is what you want for dates, because JSON has no date type and the model can only send a string. But know what lax mode accepts. It would also accept "true" for a boolean field and 1.0 for an integer. For ShipFast's schema this is harmless. If you add a field where a conversion could hide a real mistake, such as an amount in rupees, mark that one field strict with Field(strict=True) and keep the rest lax.

The general rule: be lax about how a value is written and strict about what it means. "Address Change" and address_change mean the same thing, so normalise. A date in the past means something wrong, so reject.

Check your understanding

0 of 3 answered

1.Why does ShipFast make extraction fields required but nullable, instead of optional with a default of None?

2.The model returns tracking_id: "SF48213377", correctly formatted, but the customer's message contains SF48213372. Which check catches this?

3.Where should today's date come from when validating delivery_date?