Building AI Features in Python Backends

What you will build: the ShipFast triage service


ShipFast is a parcel courier in India. Every day about 40,000 customers write to it on WhatsApp, email and in-app chat. The messages look like this: "please deliver tomorrow after 6, I'm not home", "box came wet and the charger is missing", "change address to Flat 12, Lake View, Pune 411014". Today, 60 support agents read each one, decide what it is, copy the details into the right screen and type a reply. The median message waits 3 hours before anyone acts on it.

In this course you build the service that does the first pass for them. It reads a message, decides what the customer wants, pulls out the details, sends it to the right queue and drafts a reply that an agent approves with one click. You build it in Python 3.11 with FastAPI, Pydantic v2, httpx and pytest, and you wrap the model behind one small client class you write yourself.

This first lesson shows you the finished system, so every later lesson has a place to land.

One customer message through ShipFastMessage arrivesfrom the gatewayRegex: statusquestion? Stop hereModelclassifiesthe intentModel extractsfields, if neededCode picksqueue andpriorityModeldrafts, codechecks the replyThree of six steps call a model; routing never does.
The model does the reading and writing, and every decision about where a message goes is plain Python.

The four jobs

For each message, the service does four things in order.

  1. Classify — choose one intent from a closed list: reschedule, address_change, damaged_parcel, complaint, other, or unknown.
  2. Extract — for intents that need details, pull out the tracking ID, delivery date, time window, new address and PIN code as typed fields.
  3. Route — plain Python code maps the intent and fields to a queue and a priority. The model never picks the queue.
  4. Draft — write a short, polite reply for the agent to check and send.

Two of the four jobs use a model: classify and extract. Draft uses one too. Route does not, and that is deliberate. A rule you can write in ten lines of Python is faster, cheaper and easier to test than any model call. A good AI backend is mostly ordinary backend code with a few carefully fenced model calls inside it.

The API you will end with

A customer message arrives from the messaging gateway as a small JSON body. The service accepts it at once and does the work in the background, because the three model calls together take 2 to 5 seconds.

Bash
curl -s -X POST https://api.shipfast.example/v1/messages \  -H 'Content-Type: application/json' \  -d '{"message_id": "wa-88121", "customer_id": "c-4410", "is_business": false,       "text": "SF20931847 please deliver tomorrow after 6, I am not home"}'# 202 Accepted# {"message_id": "wa-88121", "status_url": "/v1/messages/wa-88121"}

A second later, the status URL returns the finished triage:

JSON
{  "status": "done",  "result": {    "message_id": "wa-88121",    "intent": "reschedule",    "queue": "rescheduling",    "priority": "normal",    "fields": {      "tracking_id": "SF20931847",      "delivery_date": "2026-09-24",      "window_start": "18:00",      "window_end": null,      "new_address": null,      "pincode": null    },    "draft_reply": "Hi, we have requested delivery of SF20931847 on 24 September after 18:00. Thank you for letting us know.",    "prompt_versions": {"classify": "classify-v3", "extract": "extract-v2", "draft": "draft-v2"},    "cost_usd": 0.01142,    "degraded": false  }}

Look at what is in that response besides the answer. prompt_versions tells you exactly which prompts produced it, so you can trace a bad result back to a change. cost_usd tells you what this one message cost, so finance never gets a surprise. degraded says whether the service fell back to a safe path because the model provider was down. These fields are the difference between a demo and a service you can run.

Numbers to keep in mind

You will meet these numbers again and again. They are typical for ShipFast's traffic; your own will differ, and you will learn to measure them.

MeasureShipFast value
Messages per dayabout 40,000 (peak about 85 per minute)
Messages answered by a regex, no modelabout 12% ("where is SF12345678?")
Model calls per message2 or 3 (classify, maybe extract, draft)
Time for all calls, median / 95th percentile1.8 s / 4.9 s
Cost per messageabout $0.004 to $0.012
Monthly model bill at that volumeroughly $8,000 to $10,000

The last row is why cost gets its own lesson. At 40,000 messages a day, a prompt that is 300 tokens longer than it needs to be costs real money every month, and a retry loop with no limit can double the bill overnight.

The course map

Each section adds one layer to ShipFast. By the end of Section 3 you have a working extraction endpoint. By the end of Section 4 you have the full triage service shown above. Section 5 makes it safe to leave running.

SectionWhat ShipFast gains
1. Where AI fitsThe mental model: a model call is an unreliable remote function
2. Making the callLLMClient with timeouts, usage logging, cost budgets, streaming
3. Turning output into dataJSON schemas, Pydantic validation, repair, classification, extraction; Project 1: /v1/extract
4. Using it in a real serviceBackground jobs, routing, caching, retries and fallbacks; Project 2: full triage
5. Keeping it safePrompt injection defences, private-data-safe logs, per-customer quotas, prompt tests in CI

What this course does not cover

You will not train a model, fine-tune one or host one on GPUs. You will not build a chatbot that talks to customers without a human. Both are real topics, but most backend work with LLMs is what this course covers: calling a hosted model from a service, turning its reply into trusted data and keeping the service fast, cheap and safe.

Check your understanding

0 of 3 answered

1.In the ShipFast design, why does plain Python code choose the queue instead of the model?

2.Why does each triage result include prompt_versions and cost_usd?

3.The median model time for a full triage is 1.8 seconds and the 95th percentile is 4.9 seconds. What does this suggest about the API design?