Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG assistant takes 8 seconds to respond. Users bounce after 3 seconds. You cannot reduce actual generation time. How do you implement streaming + progressive rendering so users see partial answers immediately?
What you need to know
Where the 8 seconds go
Break the latency down before you design anything. A typical RAG request looks like this:
| Stage | Typical time | Can the user see something? |
|---|---|---|
| Embed the query | ~50 ms | A "Searching" status |
| Vector search | ~100 ms | — |
| Rerank | ~150 ms | Source cards, as soon as this ends |
| Prefill (model reads the prompt) | 0.5–1 s | — |
| Decode (model writes tokens) | 6–7 s | Tokens, one by one |
So something real can be on screen at about 300 ms. The long part is decoding, and streaming turns it from a wait into reading time.
Stream end to end, with no buffer in the middle
Setting stream=True on the model call is the easy part. The answer must also flow through your backend, every proxy and the browser without anyone collecting the whole body first. Server-Sent Events (SSE) is the usual transport: one long HTTP response, with small event: and data: messages flushed as they are produced.
1import json2from fastapi import FastAPI3from fastapi.responses import StreamingResponse45app = FastAPI()67def sse(event: str, data: dict) -> str:8 return f"event: {event}\ndata: {json.dumps(data)}\n\n"910@app.get("/ask")11async def ask(q: str):12 async def events():13 yield sse("status", {"text": "Searching policy documents"})14 chunks = await retrieve_and_rerank(q, top_k=5) # ~300 ms15 yield sse("sources", {"items": [c.meta for c in chunks]})16 async for token in llm.stream(build_prompt(q, chunks)): # 6-7 s17 yield sse("token", {"t": token})18 yield sse("done", {})19 return StreamingResponse(events(), media_type="text/event-stream",20 headers={"Cache-Control": "no-cache",21 "X-Accel-Buffering": "no"})The generator sends three kinds of events: a status, the sources, then tokens. retrieve_and_rerank and llm.stream stand for your own retrieval and model client. The X-Accel-Buffering: no header tells nginx not to buffer this one response. In nginx itself you can also set proxy_buffering off on the route. Other common buffers: gzip middleware, a CDN, an API gateway, and any middleware that reads the full body to log or validate it.
The progressive-rendering plan
- Status at ~200 ms — "Searching 4 sources". An honest stage message, not a spinner.
- Source cards at ~300 ms — titles and links from retrieval. Many users click one before the answer ends.
- Cut TTFT — rerank to 5 chunks instead of 20 so prefill is shorter, and keep the static system prompt first so prompt caching can reuse it.
- Stream tokens — render markdown up to the last complete block, so a half-open code fence or table does not flicker.
- Verify after — run a grounding check when the stream ends and add a "verified" badge. Moderate on a rolling window of text while streaming, not on the full body.
The measures to watch are TTFT p50 and p95 (aim for under 1 second), bounce rate, tokens per second and completion rate.
A real-life example
Scenario, numbers made up. An insurance company's policy assistant takes 8.2 seconds per answer, and analytics show 41% of users close the tab before the answer appears. The team adds streaming, but TTFT on production is still 7.9 seconds, while it is 0.9 seconds on a laptop.
Running curl -N against each hop shows the problem. The model streams fine; the backend streams fine; the nginx ingress collects the whole body. After they turn off proxy buffering on the SSE route and remove gzip from it, first token arrives at 0.9 seconds. They then add the status event and source cards at 0.3 seconds, and cut reranked chunks from 20 to 5. Total time is still about 8 seconds, but the bounce rate falls to 9%.
Follow-up questions to expect
- "Why SSE and not WebSockets?" — The traffic is one-way, server to client, so SSE is simpler: plain HTTP, automatic reconnect, and it passes through most proxies. WebSockets are worth it when the client also sends a stream, such as live voice.
- "What if the user closes the tab halfway?" — Detect the disconnect and cancel the model call, so you stop paying for tokens nobody will read.
- "How do you report an error after streaming has started?" — The HTTP status is already 200, so send an
errorevent the client knows how to show, and keep the partial text marked as incomplete.