Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Project 1: the extraction endpoint
You now have every piece of the first project: a client, a budget, a schema, a repair loop and an extraction prompt. This lesson puts them behind an HTTP endpoint that ShipFast's agent console can call. When an agent opens a message by hand, the console sends the text to /v1/extract and pre-fills the rescheduling form with the result.
This is the first time the pieces meet real HTTP semantics, and most of the work is deciding what each outcome means for the caller. A model that returned bad data twice is not a server error. A provider that is down is not the client's fault. An oversized message is. The response codes should say so.
The contract
Write the request and response models first. They are the promise you make to the console team.
1# shipfast/api.py (extraction part)2from datetime import datetime3from functools import lru_cache4from typing import Literal5from zoneinfo import ZoneInfo67from fastapi import Depends, FastAPI, HTTPException8from pydantic import BaseModel, Field910from shipfast.config import settings11from shipfast.extract import PROMPT_VERSION as EXTRACT_VERSION, extract12from shipfast.llm import LLMClient, LLMError13from shipfast.schemas import Extraction14from shipfast.structured import StructuredOutputError1516app = FastAPI(title="ShipFast support API")17IST = ZoneInfo("Asia/Kolkata")181920class ExtractRequest(BaseModel):21 message_id: str = Field(min_length=1, max_length=64)22 text: str = Field(min_length=1, max_length=2_000)232425class ExtractResponse(BaseModel):26 message_id: str27 status: Literal["ok", "needs_review"]28 fields: Extraction | None29 prompt_version: str30 cost_usd: floatstatus has two values, and the console treats them differently: ok pre-fills the form, needs_review shows an empty form with a note. fields reuses the Extraction model, so the response schema in the OpenAPI docs is exactly the validated shape. prompt_version and cost_usd travel with every response, as promised in lesson 1.
The endpoint
1# shipfast/api.py (continued)2@lru_cache3def get_llm() -> LLMClient:4 return LLMClient(settings.llm_model, timeout_s=settings.llm_timeout_s)567def today_ist():8 return datetime.now(IST).date()91011@app.post("/v1/extract", response_model=ExtractResponse)12async def extract_fields(req: ExtractRequest, llm=Depends(get_llm)) -> ExtractResponse:13 try:14 out = await extract(llm, req.text, today_ist())15 except StructuredOutputError as err:16 return ExtractResponse(message_id=req.message_id, status="needs_review", fields=None,17 prompt_version=EXTRACT_VERSION,18 cost_usd=round(sum(c.cost_usd for c in err.calls), 5))19 except LLMError:20 raise HTTPException(503, "extraction temporarily unavailable", headers={"Retry-After": "30"})21 return ExtractResponse(message_id=req.message_id, status="ok", fields=out.value,22 prompt_version=EXTRACT_VERSION,23 cost_usd=round(sum(c.cost_usd for c in out.calls), 5))Each outcome maps to one response.
| Outcome | Response | Why |
|---|---|---|
| Valid extraction | 200, status: ok | The normal case |
| Invalid after one repair | 200, status: needs_review | The request was fine and was handled; the answer is "a person must fill this in" |
| Text too long or empty | 422 | The caller sent bad input; FastAPI does this from ExtractRequest |
| Provider timeout, 429, 5xx | 503 with Retry-After | Temporary, not the caller's fault; try again later |
| Our bug (400 from provider) | 503 and an alert | Also LLMError; the log line says LLMBadRequest, which pages the team |
The second row is the one teams most often get wrong. Returning 500 when the model could not produce valid data makes dashboards show an outage when nothing is broken, and makes clients retry a request that will fail the same way. "We could not extract this" is a successful, honest answer.
The last row is a trade-off. LLMBadRequest is a bug, and some teams return 500 for it. ShipFast returns 503 so the console behaves the same way for every provider-side failure, and relies on the log level and an alert to make sure the bug is fixed. Either choice is defensible; being consistent is what matters.
How one request flows
It helps to see the whole path once, from the console's request to its response, with the pieces from earlier lessons in order.
- FastAPI validates the request —
ExtractRequestrejects empty or oversized text with a 422. No model call, no cost. - The endpoint computes today in India time — one value, used for the prompt and for validation.
extract()builds the prompt — the fixed system prompt, plus today's date and the message inside<message>tags.call_structured()calls the model with the schema — throughLLMClient, which applies the timeout, maps errors and logs usage.- Pydantic validates with context — grounding check on the tracking ID, date range, time format, PIN code.
- One repair if needed — the errors go back to the model once; a second failure becomes
needs_review. - The endpoint answers —
okwith fields,needs_review, or 503 when the provider is unavailable, always with cost and prompt version.
Nothing in this list is new. The project is the point where you see that each earlier lesson solved one step, and that the endpoint itself is mostly glue.
Running it
1pip install fastapi uvicorn anthropic pydantic httpx pytest2export ANTHROPIC_API_KEY=... # never commit this3uvicorn shipfast.api:app --reload45curl -s -X POST localhost:8000/v1/extract -H 'Content-Type: application/json' \6 -d '{"message_id": "wa-1", "text": "SF20931847 pls deliver tmrw after 6"}'1{"message_id": "wa-1", "status": "ok",2 "fields": {"tracking_id": "SF20931847", "delivery_date": "2026-09-24", "window_start": "18:00",3 "window_end": null, "new_address": null, "pincode": null},4 "prompt_version": "extract-v2", "cost_usd": 0.00451}The server log shows one line for the call: llm_call feature=extract model=claude-opus-5 in=468 out=86 ms=1240 cost_usd=0.00449 stop=end_turn.
Testing without a network
FastAPI's dependency_overrides replaces get_llm with a function that returns a FakeLLM, so tests exercise the real endpoint, the real extract(), the real validators and the real repair loop, with scripted model replies.
1# tests/test_extract_api.py2import json3from datetime import date45from fastapi.testclient import TestClient67from shipfast import api8from tests.fakes import FakeLLM910MSG = "SF12345678 please deliver tomorrow after 6"11GOOD = {"tracking_id": "SF12345678", "delivery_date": "2026-09-24", "window_start": "18:00",12 "window_end": None, "new_address": None, "pincode": None}131415def post(monkeypatch, replies: list[dict], text: str = MSG):16 fake = FakeLLM({"extract": [json.dumps(r) for r in replies]})17 monkeypatch.setattr(api, "today_ist", lambda: date(2026, 9, 23))18 api.app.dependency_overrides[api.get_llm] = lambda: fake19 try:20 return TestClient(api.app).post("/v1/extract", json={"message_id": "m1", "text": text}), fake21 finally:22 api.app.dependency_overrides.clear()232425def test_ok(monkeypatch):26 resp, fake = post(monkeypatch, [GOOD])27 assert resp.json()["status"] == "ok" and len(fake.calls) == 1282930def test_invented_id_is_repaired(monkeypatch):31 resp, fake = post(monkeypatch, [GOOD | {"tracking_id": "SF99999999"}, GOOD])32 assert resp.json()["fields"]["tracking_id"] == "SF12345678" and len(fake.calls) == 2333435def test_gives_up_after_one_repair(monkeypatch):36 bad = GOOD | {"delivery_date": "2025-01-01"}37 resp, fake = post(monkeypatch, [bad, bad])38 assert resp.json()["status"] == "needs_review" and len(fake.calls) == 2394041def test_too_long_never_calls_model(monkeypatch):42 resp, fake = post(monkeypatch, [], text="x" * 2_001)43 assert resp.status_code == 422 and fake.calls == []These four tests cover the four behaviours that matter: the happy path, one repair, giving up and rejecting bad input before spending money. They run in well under a second and cost nothing. Freezing today_ist makes date validation deterministic; without it, the tests would start failing in October.
What to watch after launch
Once the endpoint is live, four numbers tell you whether it is healthy. Put them on one dashboard, broken down by prompt_version.
- needs_review rate — about 1% is normal for ShipFast. A rise after a deploy points at the prompt or the model; a rise with no deploy points at new kinds of messages.
- Cost per call, median and 95th percentile — the median should sit near $0.0045. A growing 95th percentile means more repairs or longer messages.
- Latency, 95th percentile — about 2.5 seconds. If it climbs, check the provider's status page before your code.
- 503 rate — provider trouble. Section 4 adds retries and a fallback model so most of these never reach the console.
Check your understanding
0 of 3 answered
1.The model fails validation twice for a message. Which response should /v1/extract return?
2.Why does the test freeze today_ist?
3.What does test_too_long_never_calls_model protect against?