Building AI Features in Python Backends

Streaming responses through FastAPI


ShipFast's agents use a console. When an agent opens a message that the triage service could not draft a reply for, they can press "Draft for me". The first version waited for the whole reply and then showed it. A 90-word reply took about 3 seconds, and agents pressed the button again because nothing seemed to happen, which started a second paid call.

The model writes one token at a time, and the first tokens are ready after about half a second. Streaming sends them to the browser as they arrive. The total time does not change, but the agent sees words after 0.5 seconds instead of a blank box for 3 seconds, and stops pressing the button twice.

The same 90-word draft, two waysWait for the full reply• Blank box for about 3 s• Checked before anyone sees it• Errors travel as HTTP status• 18% of agents pressed twiceStream as server-sent events• First words after about 0.5 s• Checked when the stream ends• Errors travel as events in a 200• Repeat presses under 2%
Streaming does not make the reply faster; it moves the first visible word from three seconds to half a second.

When to stream and when not to

Stream when

  • A person reads the output as it arrives
  • The output is long: more than about 50 tokens
  • Nothing downstream needs the whole reply before acting

Do not stream when

  • Code parses the output (JSON for classification or extraction)
  • The output is short; the first token and the last arrive almost together
  • You must validate the whole reply before anyone sees it

ShipFast streams exactly one thing: the draft reply in the agent console. Classification and extraction return small JSON objects that code must validate as a whole. Streaming them would add complexity for no visible gain, because half a JSON object is useless.

There is a real trade-off here. You cannot fully validate a streamed reply until it has finished, so the agent may see words that your checks would reject. For an internal console where a trained agent reviews every draft, that is acceptable. For text going straight to a customer, it is not: generate, check, then send.

Adding stream() to LLMClient

The SDK's streaming helper gives you an async iterator of text chunks and, at the end, the final message with its usage. The client yields chunks and logs usage exactly like complete(), so streamed calls show up in the same cost reports.

Python
# shipfast/llm.py, inside class LLMClient    async def stream(self, *, feature: str, system: str, messages: list[dict],                     max_tokens: int = 600) -> AsyncIterator[str]:        start = time.perf_counter()        try:            async with self._sdk.messages.stream(                model=self.model, max_tokens=max_tokens, system=system, messages=messages,                output_config={"effort": self.effort} if self.effort else anthropic.omit,            ) as stream:                async for chunk in stream.text_stream:                    yield chunk                final = await stream.get_final_message()        except anthropic.APIError as err:            raise _translate(feature, err) from err        _log(feature, LLMResult(            text="", model=self.model, input_tokens=final.usage.input_tokens,            output_tokens=final.usage.output_tokens, stop_reason=final.stop_reason,            latency_ms=round((time.perf_counter() - start) * 1000)))

The async with block matters. If the consumer stops iterating, for example because the browser disconnected, leaving the block closes the connection to the provider. Generation stops, and you stop paying for tokens nobody will read.

Serving it as server-sent events

Server-sent events (SSE) are a simple format for a one-way stream over HTTP: lines starting with event: and data:, separated by a blank line. Browsers read them with EventSource or fetch, and they pass through most proxies and load balancers.

Python
# shipfast/api.py (streaming part)import jsonfrom functools import lru_cachefrom fastapi import Depends, FastAPIfrom fastapi.responses import StreamingResponsefrom pydantic import BaseModel, Fieldfrom shipfast.config import settingsfrom shipfast.drafts import SYSTEM as DRAFT_SYSTEM, draft_problemsfrom shipfast.llm import LLMClient, LLMErrorapp = FastAPI(title="ShipFast support API")@lru_cachedef get_llm() -> LLMClient:    return LLMClient(settings.llm_model, timeout_s=settings.llm_timeout_s)class DraftRequest(BaseModel):    text: str = Field(min_length=1, max_length=2_000)def sse(event: str, data: dict) -> str:    return f"event: {event}\ndata: {json.dumps(data)}\n\n"@app.post("/v1/drafts/stream")async def stream_draft(req: DraftRequest, llm: LLMClient = Depends(get_llm)):    async def events():        parts: list[str] = []        try:            async for chunk in llm.stream(                    feature="draft_stream", system=DRAFT_SYSTEM,                    messages=[{"role": "user", "content": f"<message>\n{req.text}\n</message>"}]):                parts.append(chunk)                yield sse("delta", {"text": chunk})        except LLMError as err:            yield sse("error", {"message": type(err).__name__})            return        yield sse("done", {"problems": draft_problems("".join(parts))})    return StreamingResponse(events(), media_type="text/event-stream",                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

DRAFT_SYSTEM and draft_problems come from shipfast/drafts.py, which Section 4 builds; for now, read them as "the draft instructions" and "a function that lists problems in a draft, such as a promised refund". @lru_cache on get_llm means one client per process, so HTTP connections to the provider are reused.

Three details in this endpoint save you from real incidents.

Errors after the first byte. Once streaming starts, the HTTP status (200) has already been sent. You cannot change it to 503. So errors travel inside the stream as an error event, and the console shows "Draft failed, write it yourself" instead of hanging.

Checks at the end. The whole draft is checked once it is complete, and the result travels in the done event. If the model wrote "we will refund you", the console shows a warning next to the draft before the agent can send it.

Proxy buffering. Many reverse proxies, nginx by default, collect a response before forwarding it, which silently turns your stream into one late block. The X-Accel-Buffering: no header disables that for nginx; check your load balancer's documentation for the equivalent, and check its idle timeout too, because a slow stream can be cut off by a 30-second limit.

Testing a stream

The endpoint only needs something with a stream() method, so a test swaps in a fake that yields fixed chunks and uses FastAPI's TestClient to read the events.

Python
from fastapi.testclient import TestClientfrom shipfast.api import app, get_llmclass FakeStreamLLM:    model = "claude-haiku-4-5"    async def stream(self, **_: object):        for part in ["Hi, ", "we have ", "requested a refund."]:            yield partdef test_stream_flags_a_refund_promise():    app.dependency_overrides[get_llm] = lambda: FakeStreamLLM()    with TestClient(app).stream("POST", "/v1/drafts/stream", json={"text": "box wet"}) as r:        body = "".join(r.iter_text())    app.dependency_overrides.clear()    assert body.count("event: delta") == 3    assert "forbidden" in body          # the done event lists the refund problem

The test checks the order of events and the final check without any network, which is exactly what can break when someone edits the endpoint.

Check your understanding

0 of 3 answered

1.Why does ShipFast not stream its classification replies?

2.A stream fails after 40 words have been sent to the agent's console. What should the endpoint do?

3.A stream works locally but arrives as one block after 3 seconds in production. What is the most likely cause?