Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Synchronous calls or background jobs
ShipFast's messaging gateway forwards every customer message to the triage service with an HTTP POST. It waits 5 seconds for a response. If it gets nothing, it assumes the delivery failed and sends the same message again, up to three times.
The first triage endpoint did all the work inside the request: classify, extract, draft, then respond. The median was 1.8 seconds, well inside 5. But the 95th percentile was 4.9 seconds, and during a busy evening the provider slowed down and the 95th percentile went past 7 seconds. The gateway timed out on about one message in eight and resent it. Each resend started a second full triage while the first was still running. The model bill for that evening was 30% higher than normal, and agents saw some messages twice with different labels.
Nothing in the model was wrong. The endpoint had the wrong shape for its caller.
The decision
Two numbers decide it: how long the caller will wait, and how long your slowest normal request takes. Use the 95th or 99th percentile, never the median.
Answer synchronously when
- The caller is a person waiting on a screen, or code that needs the answer to continue
- One short model call, with a 99th percentile well under the caller's timeout
- A failure can simply be shown and retried by the user
- Example:
/v1/extractfor the agent console, about 2.5 s at the 95th percentile
Accept and process in the background when
- The caller only needs to know the work was received
- Several calls in a row, or a long tail that can pass the caller's timeout
- The work must survive a restart or a provider outage
- Example:
/v1/messagesfrom the gateway, 2 to 3 calls per message
ShipFast's console extraction stays synchronous: an agent is waiting and one call is fast enough. Triage from the gateway becomes a background job.
Accept with 202, answer with a status URL
The endpoint does three quick things and returns: validate the body, record that the message was received, and schedule the work. That takes a few milliseconds, so the gateway never times out.
1# shipfast/api.py (triage part)2from fastapi import BackgroundTasks, Depends, HTTPException34from shipfast.schemas import InboundMessage, TriageResult5from shipfast.triage import triage67RESULTS: dict[str, TriageResult | None] = {} # stand-in for a Postgres table8910async def run_triage(msg: InboundMessage, llm, cache) -> None:11 RESULTS[msg.message_id] = await triage(llm, msg, today_ist(), cache)121314@app.post("/v1/messages", status_code=202)15async def accept_message(msg: InboundMessage, tasks: BackgroundTasks,16 llm=Depends(get_llm), cache=Depends(get_cache)) -> dict:17 if msg.message_id not in RESULTS: # idempotent: a retried webhook is a no-op18 RESULTS[msg.message_id] = None19 tasks.add_task(run_triage, msg, llm, cache)20 return {"message_id": msg.message_id, "status_url": f"/v1/messages/{msg.message_id}"}212223@app.get("/v1/messages/{message_id}")24async def get_message(message_id: str) -> dict:25 if message_id not in RESULTS:26 raise HTTPException(404, "unknown message")27 result = RESULTS[message_id]28 return {"status": "pending"} if result is None else {"status": "done", "result": result}triage() is built in the last lesson of this section; get_cache arrives with caching. For now, read triage as "the whole pipeline".
The key line is the if. The gateway's resend carries the same message_id, so it finds the entry and returns the same 202 without scheduling anything. That is the idempotency from Section 1, now at the HTTP boundary. Consumers such as the agent console either poll the status URL or, better, receive an event when the result is stored.
When BackgroundTasks is not enough
FastAPI's BackgroundTasks runs the function in the same process after the response is sent. It is simple and it works, with three limits you must accept knowingly.
- Work is lost on restart. A deploy or a crash kills every task that has not finished. The gateway already got its 202, so it will not resend.
- No retry later. If the provider is down for 20 minutes, the task fails and nothing tries again.
- No limit on concurrency. A burst of 500 messages starts 500 tasks at once, each holding a connection to the provider.
For a prototype, or for work you can afford to lose, those limits are fine. For ShipFast, where every lost message is a customer waiting, the production version stores the job in Postgres and runs separate workers.
1CREATE TABLE triage_jobs (2 message_id text PRIMARY KEY,3 payload jsonb NOT NULL,4 status text NOT NULL DEFAULT 'pending', -- pending, running, done, failed5 attempts int NOT NULL DEFAULT 0,6 result jsonb,7 updated_at timestamptz NOT NULL DEFAULT now()8);910-- Accept: a duplicate delivery inserts nothing.11INSERT INTO triage_jobs (message_id, payload) VALUES ($1, $2)12ON CONFLICT (message_id) DO NOTHING;1314-- Worker: claim the oldest job. SKIP LOCKED lets many workers run without blocking each other,15-- and a job stuck in 'running' for 2 minutes (its worker died) is picked up again.16UPDATE triage_jobs SET status = 'running', attempts = attempts + 1, updated_at = now()17WHERE message_id = (18 SELECT message_id FROM triage_jobs19 WHERE status = 'pending'20 OR (status = 'running' AND updated_at < now() - interval '2 minutes')21 ORDER BY updated_at LIMIT 122 FOR UPDATE SKIP LOCKED)23RETURNING message_id, payload;A worker loops: claim a job, run triage(), write the result and set status = 'done'. If attempts reaches 5, it sets failed and the message goes to a human queue. The endpoint's only job becomes the INSERT. A queue library (Celery, RQ, arq) or a managed queue gives you the same shape; the table version is shown because every backend team already runs Postgres and it makes the idempotency explicit.
How many workers?
Little's law gives the answer in one line: work in progress equals arrival rate times time in the system. ShipFast's peak is about 85 messages a minute, or 1.4 a second, and a triage takes 1.8 seconds on median. So on average 1.4 × 1.8 = about 2.5 triages are running at once. Allow for the tail (4.9 s at the 95th percentile) and for a 3× burst: 1.4 × 3 × 4.9 is about 21. ShipFast runs two workers with 16 concurrent tasks each, 32 in total, and still sits far below the provider's request limit of a few thousand per minute.
Workers are cheap, because they spend almost all their time waiting on the network. An async worker with 16 tasks uses one CPU core lightly. The limit you hit first is usually the provider's rate limit, which Section 5 deals with.
Check your understanding
0 of 3 answered
1.The gateway times out after 5 seconds. Triage has a median of 1.8 s and a 95th percentile of 4.9 s. Why is a synchronous endpoint risky?
2.A deploy restarts the service while 40 triages are running as FastAPI BackgroundTasks. What happens to them?
3.Arrival rate is 2 messages per second and each triage takes 3 seconds. About how many triages are in progress at once on average?