Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG assistant takes 8 seconds to respond. Users bounce after 3 seconds. You cannot reduce actual generation time. How do you implement streaming + progressive rendering so users see partial answers immediately?


What the user sees, and when200 ms —"Searching4 sources"300 ms —sourcecards appearabout 900 ms— firsttoken streams8 s —answer endsThen —"verified" badgeOne buffering proxy anywhere in the path pushes everything to the 8 s mark.
Total time stays at 8 seconds; what changes is that the blank screen shrinks from 8 seconds to 200 milliseconds.

What you need to know

Where the 8 seconds go

Break the latency down before you design anything. A typical RAG request looks like this:

StageTypical timeCan the user see something?
Embed the query~50 msA "Searching" status
Vector search~100 ms—
Rerank~150 msSource cards, as soon as this ends
Prefill (model reads the prompt)0.5–1 s—
Decode (model writes tokens)6–7 sTokens, one by one

So something real can be on screen at about 300 ms. The long part is decoding, and streaming turns it from a wait into reading time.

Stream end to end, with no buffer in the middle

Setting stream=True on the model call is the easy part. The answer must also flow through your backend, every proxy and the browser without anyone collecting the whole body first. Server-Sent Events (SSE) is the usual transport: one long HTTP response, with small event: and data: messages flushed as they are produced.

Python
import jsonfrom fastapi import FastAPIfrom fastapi.responses import StreamingResponseapp = FastAPI()def sse(event: str, data: dict) -> str:    return f"event: {event}\ndata: {json.dumps(data)}\n\n"@app.get("/ask")async def ask(q: str):    async def events():        yield sse("status", {"text": "Searching policy documents"})        chunks = await retrieve_and_rerank(q, top_k=5)          # ~300 ms        yield sse("sources", {"items": [c.meta for c in chunks]})        async for token in llm.stream(build_prompt(q, chunks)):  # 6-7 s            yield sse("token", {"t": token})        yield sse("done", {})    return StreamingResponse(events(), media_type="text/event-stream",                             headers={"Cache-Control": "no-cache",                                      "X-Accel-Buffering": "no"})

The generator sends three kinds of events: a status, the sources, then tokens. retrieve_and_rerank and llm.stream stand for your own retrieval and model client. The X-Accel-Buffering: no header tells nginx not to buffer this one response. In nginx itself you can also set proxy_buffering off on the route. Other common buffers: gzip middleware, a CDN, an API gateway, and any middleware that reads the full body to log or validate it.

The progressive-rendering plan

  1. Status at ~200 ms — "Searching 4 sources". An honest stage message, not a spinner.
  2. Source cards at ~300 ms — titles and links from retrieval. Many users click one before the answer ends.
  3. Cut TTFT — rerank to 5 chunks instead of 20 so prefill is shorter, and keep the static system prompt first so prompt caching can reuse it.
  4. Stream tokens — render markdown up to the last complete block, so a half-open code fence or table does not flicker.
  5. Verify after — run a grounding check when the stream ends and add a "verified" badge. Moderate on a rolling window of text while streaming, not on the full body.

The measures to watch are TTFT p50 and p95 (aim for under 1 second), bounce rate, tokens per second and completion rate.

A real-life example

Scenario, numbers made up. An insurance company's policy assistant takes 8.2 seconds per answer, and analytics show 41% of users close the tab before the answer appears. The team adds streaming, but TTFT on production is still 7.9 seconds, while it is 0.9 seconds on a laptop.

Running curl -N against each hop shows the problem. The model streams fine; the backend streams fine; the nginx ingress collects the whole body. After they turn off proxy buffering on the SSE route and remove gzip from it, first token arrives at 0.9 seconds. They then add the status event and source cards at 0.3 seconds, and cut reranked chunks from 20 to 5. Total time is still about 8 seconds, but the bounce rate falls to 9%.

Follow-up questions to expect

  • "Why SSE and not WebSockets?" — The traffic is one-way, server to client, so SSE is simpler: plain HTTP, automatic reconnect, and it passes through most proxies. WebSockets are worth it when the client also sends a stream, such as live voice.
  • "What if the user closes the tab halfway?" — Detect the disconnect and cancel the model call, so you stop paying for tokens nobody will read.
  • "How do you report an error after streaming has started?" — The HTTP status is already 200, so send an error event the client knows how to show, and keep the partial text marked as incomplete.