Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do streaming responses improve user experience in real-time systems?
What you need to know
The numbers
End-to-end time is roughly:
e2e = TTFT + output_tokens × TPOT = 0.4 s + 400 × 0.025 s = 10.4 sTTFT is time to first token (queueing plus processing the prompt). TPOT is time per output token. Without streaming, the user stares at a spinner for 10.4 seconds. With streaming, text appears at 0.4 seconds and then flows at 40 tokens a second — faster than most people read.
Blocking response
- Nothing visible for 10.4 s
- Users click "send" again, doubling cost
- Full output can be validated before showing
- Simple to implement
Streaming response
- First words at 0.4 s
- User can stop a wrong answer early
- Output checks must work on partial text
- Needs SSE plumbing end to end
How it is implemented
Server-Sent Events (SSE) is a simple HTTP format: the server keeps the response open and writes data: ... lines. Model APIs use it, and your backend relays it.
1import asyncio, json2from fastapi import FastAPI, Request3from fastapi.responses import StreamingResponse45app = FastAPI()67async def fake_model_stream(question: str):8 for word in ["Your", " refund", " was", " raised", " on", " 12", " Oct."]:9 await asyncio.sleep(0.05)10 yield word1112@app.get("/chat")13async def chat(q: str, request: Request):14 async def events():15 async for token in fake_model_stream(q):16 if await request.is_disconnected(): # user closed the tab17 break # stop paying for tokens18 yield f"data: {json.dumps({'delta': token})}\n\n"19 yield "data: [DONE]\n\n"20 return StreamingResponse(events(), media_type="text/event-stream",21 headers={"Cache-Control": "no-cache",22 "X-Accel-Buffering": "no"})In a real service, fake_model_stream is the model client called with stream=True. The disconnect check lets you cancel the upstream request, so an abandoned answer stops using GPU time. X-Accel-Buffering: no tells nginx not to buffer.
Problems streaming creates
- Buffering proxies. A load balancer or nginx that buffers the response makes streaming look broken. Test through the real network path.
- Output guardrails. You cannot check the whole answer before the user sees it. Check sentence by sentence, or buffer and validate structured output fully.
- Mid-stream errors. The HTTP status is already 200. Send an error event the client understands.
- Metrics. Measure TTFT and inter-token latency separately; an average end-to-end time hides a bad TTFT.
A real-life example
A government service chatbot runs in 12 languages, and many users are on slow mobile networks. Answers in Indian languages are often 500–700 tokens because these scripts take more tokens per word. Before streaming, answers took 14 seconds on average; analytics showed 22% of users sent the question again before the first answer arrived, which doubled load at the worst moments.
After streaming, first words appear in about 0.8 seconds. Duplicate sends fall to 3%, which cuts GPU load by almost a fifth at peak. When users close the app mid-answer (about 9% of answers), the backend cancels generation, saving those tokens too. The one bug at launch: a CDN rule buffered text/event-stream, so the first week's mobile users still saw the whole answer at once until it was excluded.
Follow-up questions to expect
- "SSE or WebSockets?" — SSE for one-way token streams: plain HTTP, works through most proxies, auto-reconnects. WebSockets when you need two-way real-time messages, such as voice.
- "How do you stream JSON?" — Usually you do not show it until complete; or stream partial objects with a parser that tolerates incomplete JSON.
- "Does streaming change cost?" — Not per token, but cancelling abandoned answers and preventing duplicate sends reduces total tokens.