Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement streaming responses from an LLM API.
What you need to know
A model writes one token at a time. Without streaming, the server waits for the last token before sending anything, so a 600-token answer at 50 tokens per second shows a blank screen for 12 seconds. With streaming, the first words appear after the time to first token (TTFT), usually well under a second.
Server-sent events (SSE) is the simplest way to push a stream to a browser. It is a normal HTTP response with Content-Type: text/event-stream that stays open; each event is a line starting with data: followed by a blank line:
data: {"delta": "Refunds"}data: {"delta": " take 5 days."}data: [DONE]Three production details catch people out:
- Buffering. Nginx and some load balancers collect the whole response before forwarding it. The stream "works" but arrives all at once. The header
X-Accel-Buffering: noturns that off for nginx. - Errors after the first byte. The status line (200) has already been sent, so a failure halfway through cannot become a 500. It must be an in-band error event the client understands.
- Cancellation. If the user closes the tab, stop generating; otherwise you pay for tokens nobody reads.
The model call, with the Anthropic Python SDK:
1from collections.abc import Iterator2import logging3from anthropic import Anthropic45client = Anthropic()6log = logging.getLogger("llm")78def stream_answer(prompt: str) -> Iterator[str]:9 """Yield text deltas as the model produces them; log usage at the end."""10 with client.messages.stream(11 model="claude-opus-5", max_tokens=16000,12 messages=[{"role": "user", "content": prompt}],13 ) as stream:14 yield from stream.text_stream15 final = stream.get_final_message()16 log.info("usage in=%s out=%s stop=%s", final.usage.input_tokens,17 final.usage.output_tokens, final.stop_reason)The HTTP endpoint, with FastAPI:
1import json2from fastapi import FastAPI3from fastapi.responses import StreamingResponse4from pydantic import BaseModel56app = FastAPI()78class ChatRequest(BaseModel):9 prompt: str1011def sse(payload: dict | str) -> str:12 return f"data: {payload if isinstance(payload, str) else json.dumps(payload)}\n\n"1314@app.post("/chat")15def chat(req: ChatRequest) -> StreamingResponse:16 def events() -> Iterator[str]:17 try:18 for delta in stream_answer(req.prompt):19 yield sse({"delta": delta})20 except Exception as exc: # headers already sent: report in-band21 yield sse({"error": type(exc).__name__})22 yield sse("[DONE]")23 return StreamingResponse(events(), media_type="text/event-stream",24 headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})The tricky parts:
with client.messages.stream(...)closes the HTTP connection to the provider when the block exits. When the browser disconnects, the web server stops iteratingevents(); once the generator is closed (explicitly or by garbage collection), thewithblock exits and generation stops. In an async app, useAsyncAnthropicand checkawait request.is_disconnected()to stop even sooner.yield from stream.text_streamgives only text deltas; usage andstop_reasonexist only once the stream is finished, so cost logging happens after the loop.- The error event carries only the exception type, not its message, which can contain internal details.
- The request body is a Pydantic model. A bare
prompt: strparameter on a POST route becomes a query-string parameter in FastAPI, which ends up in access logs.
Complexity: total time is unchanged — the same tokens are generated. Server memory is O(1 chunk) instead of O(whole response). Each event adds a few bytes of framing.
A real-life example
A fake stream_answer that yields three chunks and then fails lets us see the exact bytes:
1from fastapi.testclient import TestClient23def fake_stream(prompt: str):4 yield "Refunds"5 yield " take 5 days."6 raise TimeoutError("provider stalled")78stream_answer = fake_stream9print(TestClient(app).post("/chat", json={"prompt": "refund time?"}).text)data: {"delta": "Refunds"}data: {"delta": " take 5 days."}data: {"error": "TimeoutError"}data: [DONE]| moment | server state | client sees |
|---|---|---|
| chunk 1 | status 200 and headers sent | "Refunds" |
| chunk 2 | still streaming | "Refunds take 5 days." |
| failure | cannot change the status any more | an error event: show "answer cut off, retry?" |
| end | generator finished | [DONE]: close the connection |
Every consumer chat product — a bank's assistant, a food app's support chat — streams this way, because a 10-second blank screen makes users hit refresh and pay for the answer twice.
Follow-up questions to expect
- "How do you stream structured JSON?" — Partial JSON does not parse. Either stream only a text field and send the structure at the end, or use a tolerant partial-JSON parser on the client.
- "The answer stops mid-sentence — why?" — Check
stop_reason;max_tokensmeans it hit the output limit. Surface it to the client instead of pretending the answer is complete. - "SSE or WebSockets?" — SSE for one-way model output: plain HTTP, automatic reconnect, works through proxies. WebSockets when the client also sends a stream, such as live audio.