Building AI Features in Python Backends

Synchronous calls or background jobs


ShipFast's messaging gateway forwards every customer message to the triage service with an HTTP POST. It waits 5 seconds for a response. If it gets nothing, it assumes the delivery failed and sends the same message again, up to three times.

The first triage endpoint did all the work inside the request: classify, extract, draft, then respond. The median was 1.8 seconds, well inside 5. But the 95th percentile was 4.9 seconds, and during a busy evening the provider slowed down and the 95th percentile went past 7 seconds. The gateway timed out on about one message in eight and resent it. Each resend started a second full triage while the first was still running. The model bill for that evening was 30% higher than normal, and agents saw some messages twice with different labels.

Nothing in the model was wrong. The endpoint had the wrong shape for its caller.

Accept in milliseconds, work in the backgroundGateway POSTs,waits 5 s at mostINSERT ...ON CONFLICTDO NOTHING202 and astatus URL,in about 15 msWorker claimswith SKIP LOCKEDtriage() runs,result storedA resend with the same message_id inserts nothing.
The caller's timeout is measured against your slowest normal request, so multi-call work moves out of the request.

The decision

Two numbers decide it: how long the caller will wait, and how long your slowest normal request takes. Use the 95th or 99th percentile, never the median.

Answer synchronously when

  • The caller is a person waiting on a screen, or code that needs the answer to continue
  • One short model call, with a 99th percentile well under the caller's timeout
  • A failure can simply be shown and retried by the user
  • Example: /v1/extract for the agent console, about 2.5 s at the 95th percentile

Accept and process in the background when

  • The caller only needs to know the work was received
  • Several calls in a row, or a long tail that can pass the caller's timeout
  • The work must survive a restart or a provider outage
  • Example: /v1/messages from the gateway, 2 to 3 calls per message

ShipFast's console extraction stays synchronous: an agent is waiting and one call is fast enough. Triage from the gateway becomes a background job.

Accept with 202, answer with a status URL

The endpoint does three quick things and returns: validate the body, record that the message was received, and schedule the work. That takes a few milliseconds, so the gateway never times out.

Python
# shipfast/api.py (triage part)from fastapi import BackgroundTasks, Depends, HTTPExceptionfrom shipfast.schemas import InboundMessage, TriageResultfrom shipfast.triage import triageRESULTS: dict[str, TriageResult | None] = {}   # stand-in for a Postgres tableasync def run_triage(msg: InboundMessage, llm, cache) -> None:    RESULTS[msg.message_id] = await triage(llm, msg, today_ist(), cache)@app.post("/v1/messages", status_code=202)async def accept_message(msg: InboundMessage, tasks: BackgroundTasks,                         llm=Depends(get_llm), cache=Depends(get_cache)) -> dict:    if msg.message_id not in RESULTS:            # idempotent: a retried webhook is a no-op        RESULTS[msg.message_id] = None        tasks.add_task(run_triage, msg, llm, cache)    return {"message_id": msg.message_id, "status_url": f"/v1/messages/{msg.message_id}"}@app.get("/v1/messages/{message_id}")async def get_message(message_id: str) -> dict:    if message_id not in RESULTS:        raise HTTPException(404, "unknown message")    result = RESULTS[message_id]    return {"status": "pending"} if result is None else {"status": "done", "result": result}

triage() is built in the last lesson of this section; get_cache arrives with caching. For now, read triage as "the whole pipeline".

The key line is the if. The gateway's resend carries the same message_id, so it finds the entry and returns the same 202 without scheduling anything. That is the idempotency from Section 1, now at the HTTP boundary. Consumers such as the agent console either poll the status URL or, better, receive an event when the result is stored.

When BackgroundTasks is not enough

FastAPI's BackgroundTasks runs the function in the same process after the response is sent. It is simple and it works, with three limits you must accept knowingly.

  • Work is lost on restart. A deploy or a crash kills every task that has not finished. The gateway already got its 202, so it will not resend.
  • No retry later. If the provider is down for 20 minutes, the task fails and nothing tries again.
  • No limit on concurrency. A burst of 500 messages starts 500 tasks at once, each holding a connection to the provider.

For a prototype, or for work you can afford to lose, those limits are fine. For ShipFast, where every lost message is a customer waiting, the production version stores the job in Postgres and runs separate workers.

SQL
CREATE TABLE triage_jobs (    message_id  text PRIMARY KEY,    payload     jsonb NOT NULL,    status      text NOT NULL DEFAULT 'pending',   -- pending, running, done, failed    attempts    int NOT NULL DEFAULT 0,    result      jsonb,    updated_at  timestamptz NOT NULL DEFAULT now());-- Accept: a duplicate delivery inserts nothing.INSERT INTO triage_jobs (message_id, payload) VALUES ($1, $2)ON CONFLICT (message_id) DO NOTHING;-- Worker: claim the oldest job. SKIP LOCKED lets many workers run without blocking each other,-- and a job stuck in 'running' for 2 minutes (its worker died) is picked up again.UPDATE triage_jobs SET status = 'running', attempts = attempts + 1, updated_at = now()WHERE message_id = (    SELECT message_id FROM triage_jobs    WHERE status = 'pending'       OR (status = 'running' AND updated_at < now() - interval '2 minutes')    ORDER BY updated_at LIMIT 1    FOR UPDATE SKIP LOCKED)RETURNING message_id, payload;

A worker loops: claim a job, run triage(), write the result and set status = 'done'. If attempts reaches 5, it sets failed and the message goes to a human queue. The endpoint's only job becomes the INSERT. A queue library (Celery, RQ, arq) or a managed queue gives you the same shape; the table version is shown because every backend team already runs Postgres and it makes the idempotency explicit.

How many workers?

Little's law gives the answer in one line: work in progress equals arrival rate times time in the system. ShipFast's peak is about 85 messages a minute, or 1.4 a second, and a triage takes 1.8 seconds on median. So on average 1.4 × 1.8 = about 2.5 triages are running at once. Allow for the tail (4.9 s at the 95th percentile) and for a 3× burst: 1.4 × 3 × 4.9 is about 21. ShipFast runs two workers with 16 concurrent tasks each, 32 in total, and still sits far below the provider's request limit of a few thousand per minute.

Workers are cheap, because they spend almost all their time waiting on the network. An async worker with 16 tasks uses one CPU core lightly. The limit you hit first is usually the provider's rate limit, which Section 5 deals with.

Check your understanding

0 of 3 answered

1.The gateway times out after 5 seconds. Triage has a median of 1.8 s and a 95th percentile of 4.9 s. Why is a synchronous endpoint risky?

2.A deploy restarts the service while 40 triages are running as FastAPI BackgroundTasks. What happens to them?

3.Arrival rate is 2 messages per second and each triage takes 3 seconds. About how many triages are in progress at once on average?