Building AI Features in Python Backends

Project 1: the extraction endpoint


You now have every piece of the first project: a client, a budget, a schema, a repair loop and an extraction prompt. This lesson puts them behind an HTTP endpoint that ShipFast's agent console can call. When an agent opens a message by hand, the console sends the text to /v1/extract and pre-fills the rescheduling form with the result.

This is the first time the pieces meet real HTTP semantics, and most of the work is deciding what each outcome means for the caller. A model that returned bad data twice is not a server error. A provider that is down is not the client's fault. An oversized message is. The response codes should say so.

Every outcome of POST /v1/extract422 — bad input, no model call200 ok — fields, cost, prompt version200 needs_review — invalid after repair503, Retry-After — provider down
A model that could not produce valid data is a handled request with an honest answer, not a server error.

The contract

Write the request and response models first. They are the promise you make to the console team.

Python
# shipfast/api.py (extraction part)from datetime import datetimefrom functools import lru_cachefrom typing import Literalfrom zoneinfo import ZoneInfofrom fastapi import Depends, FastAPI, HTTPExceptionfrom pydantic import BaseModel, Fieldfrom shipfast.config import settingsfrom shipfast.extract import PROMPT_VERSION as EXTRACT_VERSION, extractfrom shipfast.llm import LLMClient, LLMErrorfrom shipfast.schemas import Extractionfrom shipfast.structured import StructuredOutputErrorapp = FastAPI(title="ShipFast support API")IST = ZoneInfo("Asia/Kolkata")class ExtractRequest(BaseModel):    message_id: str = Field(min_length=1, max_length=64)    text: str = Field(min_length=1, max_length=2_000)class ExtractResponse(BaseModel):    message_id: str    status: Literal["ok", "needs_review"]    fields: Extraction | None    prompt_version: str    cost_usd: float

status has two values, and the console treats them differently: ok pre-fills the form, needs_review shows an empty form with a note. fields reuses the Extraction model, so the response schema in the OpenAPI docs is exactly the validated shape. prompt_version and cost_usd travel with every response, as promised in lesson 1.

The endpoint

Python
# shipfast/api.py (continued)@lru_cachedef get_llm() -> LLMClient:    return LLMClient(settings.llm_model, timeout_s=settings.llm_timeout_s)def today_ist():    return datetime.now(IST).date()@app.post("/v1/extract", response_model=ExtractResponse)async def extract_fields(req: ExtractRequest, llm=Depends(get_llm)) -> ExtractResponse:    try:        out = await extract(llm, req.text, today_ist())    except StructuredOutputError as err:        return ExtractResponse(message_id=req.message_id, status="needs_review", fields=None,                               prompt_version=EXTRACT_VERSION,                               cost_usd=round(sum(c.cost_usd for c in err.calls), 5))    except LLMError:        raise HTTPException(503, "extraction temporarily unavailable", headers={"Retry-After": "30"})    return ExtractResponse(message_id=req.message_id, status="ok", fields=out.value,                           prompt_version=EXTRACT_VERSION,                           cost_usd=round(sum(c.cost_usd for c in out.calls), 5))

Each outcome maps to one response.

OutcomeResponseWhy
Valid extraction200, status: okThe normal case
Invalid after one repair200, status: needs_reviewThe request was fine and was handled; the answer is "a person must fill this in"
Text too long or empty422The caller sent bad input; FastAPI does this from ExtractRequest
Provider timeout, 429, 5xx503 with Retry-AfterTemporary, not the caller's fault; try again later
Our bug (400 from provider)503 and an alertAlso LLMError; the log line says LLMBadRequest, which pages the team

The second row is the one teams most often get wrong. Returning 500 when the model could not produce valid data makes dashboards show an outage when nothing is broken, and makes clients retry a request that will fail the same way. "We could not extract this" is a successful, honest answer.

The last row is a trade-off. LLMBadRequest is a bug, and some teams return 500 for it. ShipFast returns 503 so the console behaves the same way for every provider-side failure, and relies on the log level and an alert to make sure the bug is fixed. Either choice is defensible; being consistent is what matters.

How one request flows

It helps to see the whole path once, from the console's request to its response, with the pieces from earlier lessons in order.

  1. FastAPI validates the request — ExtractRequest rejects empty or oversized text with a 422. No model call, no cost.
  2. The endpoint computes today in India time — one value, used for the prompt and for validation.
  3. extract() builds the prompt — the fixed system prompt, plus today's date and the message inside <message> tags.
  4. call_structured() calls the model with the schema — through LLMClient, which applies the timeout, maps errors and logs usage.
  5. Pydantic validates with context — grounding check on the tracking ID, date range, time format, PIN code.
  6. One repair if needed — the errors go back to the model once; a second failure becomes needs_review.
  7. The endpoint answers — ok with fields, needs_review, or 503 when the provider is unavailable, always with cost and prompt version.

Nothing in this list is new. The project is the point where you see that each earlier lesson solved one step, and that the endpoint itself is mostly glue.

Running it

Bash
pip install fastapi uvicorn anthropic pydantic httpx pytestexport ANTHROPIC_API_KEY=...          # never commit thisuvicorn shipfast.api:app --reloadcurl -s -X POST localhost:8000/v1/extract -H 'Content-Type: application/json' \  -d '{"message_id": "wa-1", "text": "SF20931847 pls deliver tmrw after 6"}'
JSON
{"message_id": "wa-1", "status": "ok", "fields": {"tracking_id": "SF20931847", "delivery_date": "2026-09-24", "window_start": "18:00",            "window_end": null, "new_address": null, "pincode": null}, "prompt_version": "extract-v2", "cost_usd": 0.00451}

The server log shows one line for the call: llm_call feature=extract model=claude-opus-5 in=468 out=86 ms=1240 cost_usd=0.00449 stop=end_turn.

Testing without a network

FastAPI's dependency_overrides replaces get_llm with a function that returns a FakeLLM, so tests exercise the real endpoint, the real extract(), the real validators and the real repair loop, with scripted model replies.

Python
# tests/test_extract_api.pyimport jsonfrom datetime import datefrom fastapi.testclient import TestClientfrom shipfast import apifrom tests.fakes import FakeLLMMSG = "SF12345678 please deliver tomorrow after 6"GOOD = {"tracking_id": "SF12345678", "delivery_date": "2026-09-24", "window_start": "18:00",        "window_end": None, "new_address": None, "pincode": None}def post(monkeypatch, replies: list[dict], text: str = MSG):    fake = FakeLLM({"extract": [json.dumps(r) for r in replies]})    monkeypatch.setattr(api, "today_ist", lambda: date(2026, 9, 23))    api.app.dependency_overrides[api.get_llm] = lambda: fake    try:        return TestClient(api.app).post("/v1/extract", json={"message_id": "m1", "text": text}), fake    finally:        api.app.dependency_overrides.clear()def test_ok(monkeypatch):    resp, fake = post(monkeypatch, [GOOD])    assert resp.json()["status"] == "ok" and len(fake.calls) == 1def test_invented_id_is_repaired(monkeypatch):    resp, fake = post(monkeypatch, [GOOD | {"tracking_id": "SF99999999"}, GOOD])    assert resp.json()["fields"]["tracking_id"] == "SF12345678" and len(fake.calls) == 2def test_gives_up_after_one_repair(monkeypatch):    bad = GOOD | {"delivery_date": "2025-01-01"}    resp, fake = post(monkeypatch, [bad, bad])    assert resp.json()["status"] == "needs_review" and len(fake.calls) == 2def test_too_long_never_calls_model(monkeypatch):    resp, fake = post(monkeypatch, [], text="x" * 2_001)    assert resp.status_code == 422 and fake.calls == []

These four tests cover the four behaviours that matter: the happy path, one repair, giving up and rejecting bad input before spending money. They run in well under a second and cost nothing. Freezing today_ist makes date validation deterministic; without it, the tests would start failing in October.

What to watch after launch

Once the endpoint is live, four numbers tell you whether it is healthy. Put them on one dashboard, broken down by prompt_version.

  • needs_review rate — about 1% is normal for ShipFast. A rise after a deploy points at the prompt or the model; a rise with no deploy points at new kinds of messages.
  • Cost per call, median and 95th percentile — the median should sit near $0.0045. A growing 95th percentile means more repairs or longer messages.
  • Latency, 95th percentile — about 2.5 seconds. If it climbs, check the provider's status page before your code.
  • 503 rate — provider trouble. Section 4 adds retries and a fallback model so most of these never reach the console.

Check your understanding

0 of 3 answered

1.The model fails validation twice for a message. Which response should /v1/extract return?

2.Why does the test freeze today_ist?

3.What does test_too_long_never_calls_model protect against?