Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement streaming responses from an LLM API.


What you need to know

A model writes one token at a time. Without streaming, the server waits for the last token before sending anything, so a 600-token answer at 50 tokens per second shows a blank screen for 12 seconds. With streaming, the first words appear after the time to first token (TTFT), usually well under a second.

Server-sent events (SSE) is the simplest way to push a stream to a browser. It is a normal HTTP response with Content-Type: text/event-stream that stays open; each event is a line starting with data: followed by a blank line:

Text
data: {"delta": "Refunds"}data: {"delta": " take 5 days."}data: [DONE]

Three production details catch people out:

  • Buffering. Nginx and some load balancers collect the whole response before forwarding it. The stream "works" but arrives all at once. The header X-Accel-Buffering: no turns that off for nginx.
  • Errors after the first byte. The status line (200) has already been sent, so a failure halfway through cannot become a 500. It must be an in-band error event the client understands.
  • Cancellation. If the user closes the tab, stop generating; otherwise you pay for tokens nobody reads.

The model call, with the Anthropic Python SDK:

Python
from collections.abc import Iteratorimport loggingfrom anthropic import Anthropicclient = Anthropic()log = logging.getLogger("llm")def stream_answer(prompt: str) -> Iterator[str]:    """Yield text deltas as the model produces them; log usage at the end."""    with client.messages.stream(        model="claude-opus-5", max_tokens=16000,        messages=[{"role": "user", "content": prompt}],    ) as stream:        yield from stream.text_stream        final = stream.get_final_message()        log.info("usage in=%s out=%s stop=%s", final.usage.input_tokens,                 final.usage.output_tokens, final.stop_reason)

The HTTP endpoint, with FastAPI:

Python
import jsonfrom fastapi import FastAPIfrom fastapi.responses import StreamingResponsefrom pydantic import BaseModelapp = FastAPI()class ChatRequest(BaseModel):    prompt: strdef sse(payload: dict | str) -> str:    return f"data: {payload if isinstance(payload, str) else json.dumps(payload)}\n\n"@app.post("/chat")def chat(req: ChatRequest) -> StreamingResponse:    def events() -> Iterator[str]:        try:            for delta in stream_answer(req.prompt):                yield sse({"delta": delta})        except Exception as exc:                   # headers already sent: report in-band            yield sse({"error": type(exc).__name__})        yield sse("[DONE]")    return StreamingResponse(events(), media_type="text/event-stream",                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

The tricky parts:

  • with client.messages.stream(...) closes the HTTP connection to the provider when the block exits. When the browser disconnects, the web server stops iterating events(); once the generator is closed (explicitly or by garbage collection), the with block exits and generation stops. In an async app, use AsyncAnthropic and check await request.is_disconnected() to stop even sooner.
  • yield from stream.text_stream gives only text deltas; usage and stop_reason exist only once the stream is finished, so cost logging happens after the loop.
  • The error event carries only the exception type, not its message, which can contain internal details.
  • The request body is a Pydantic model. A bare prompt: str parameter on a POST route becomes a query-string parameter in FastAPI, which ends up in access logs.

Complexity: total time is unchanged — the same tokens are generated. Server memory is O(1 chunk) instead of O(whole response). Each event adds a few bytes of framing.

A real-life example

A fake stream_answer that yields three chunks and then fails lets us see the exact bytes:

Python
from fastapi.testclient import TestClientdef fake_stream(prompt: str):    yield "Refunds"    yield " take 5 days."    raise TimeoutError("provider stalled")stream_answer = fake_streamprint(TestClient(app).post("/chat", json={"prompt": "refund time?"}).text)
Text
data: {"delta": "Refunds"}data: {"delta": " take 5 days."}data: {"error": "TimeoutError"}data: [DONE]
momentserver stateclient sees
chunk 1status 200 and headers sent"Refunds"
chunk 2still streaming"Refunds take 5 days."
failurecannot change the status any morean error event: show "answer cut off, retry?"
endgenerator finished[DONE]: close the connection

Every consumer chat product — a bank's assistant, a food app's support chat — streams this way, because a 10-second blank screen makes users hit refresh and pay for the answer twice.

Follow-up questions to expect

  • "How do you stream structured JSON?" — Partial JSON does not parse. Either stream only a text field and send the structure at the end, or use a tolerant partial-JSON parser on the client.
  • "The answer stops mid-sentence — why?" — Check stop_reason; max_tokens means it hit the output limit. Surface it to the client instead of pretending the answer is complete.
  • "SSE or WebSockets?" — SSE for one-way model output: plain HTTP, automatic reconnect, works through proxies. WebSockets when the client also sends a stream, such as live audio.