Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 8: Streaming Response Integration


Scenario: a LangChain RAG app needs to stream answers to a web client, and the retrieval step sits in front of generation. How do you build it?

What the browser receives, and whenGET /ask opensan SSE streamsources eventat about 400 msfirst tokenat about 1.1 stoken eventsas generateddone eventX-Accel-Buffering: no, or nginx holds it all.
Total time barely changed; showing sources and the first words early is what made the assistant feel fast.

What you need to know

chain.stream() yields the output of the last step. In a RAG chain, that means you see nothing during retrieval, then tokens. astream_events yields events from every step, each labelled with its type, so the client can show progress: "found 4 sources", then the answer as it is written.

The event loop

Python
async def events(question: str):    async for ev in chain.astream_events({"question": question}, version="v2"):        kind = ev["event"]        if kind == "on_retriever_end":            docs = ev["data"]["output"]            yield sse("sources", [d.metadata.get("source") for d in docs])        elif kind == "on_chat_model_stream":            yield sse("token", ev["data"]["chunk"].content)    yield sse("done", {})

Sources arrive the moment retrieval ends, often a second before the first token, so the user sees progress immediately. For LangGraph agents, graph.astream(..., stream_mode="messages") streams model tokens with the node that produced them, which is the equivalent pattern.

The FastAPI side

Python
from fastapi.responses import StreamingResponse@app.get("/ask")async def ask(q: str):    return StreamingResponse(events(q), media_type="text/event-stream",                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

X-Accel-Buffering: no tells nginx not to buffer the response. Without it, or the equivalent setting on your load balancer, the proxy collects the whole stream and releases it at the end. Everything works on a laptop and "doesn't stream" in production.

Failure modes to handle explicitly

FailureWhat to do
Client disconnects mid-streamCatch the cancellation, save the partial answer and its cost, stop generation
Error after tokens were sentSend an error event; the HTTP status is already 200 and cannot change
Moderation neededCheck each sentence before release, not the final buffer
Slow first tokenSend a "searching" event at once so the UI isn't blank

SSE or WebSockets?

Server-sent events

  • One-way, server to client
  • Plain HTTP; works through most proxies
  • Automatic reconnect in browsers
  • Right for chat answers

WebSockets

  • Two-way, both directions
  • Needs upgrade support at every hop
  • More connection state to manage
  • Right for live collaboration or voice

A real-life example

Scenario (illustrative numbers). A university's admissions assistant shows a spinner for 5 to 6 seconds, then the full answer. Engineers added streaming, which works locally, but in production the answer still appears all at once.

The cause is nginx buffering the response. Adding X-Accel-Buffering: no fixes it. Switching from stream to astream_events lets them send sources after about 400 ms and the first token after about 1.1 s. They also save partial answers on disconnect: about 7% of users close the tab mid-answer, and those conversations are now resumable. In a survey, "the assistant feels fast" rises from 41% to 78%, though total generation time barely changed.

Follow-up questions to expect

  • "Why not WebSockets?" — SSE is simpler and enough for one-way streaming; choose WebSockets when the client must send data during the stream.
  • "How do you stream agent tool progress?" — Emit events for tool start and end ("checking your order...") from astream_events or LangGraph's stream modes.
  • "How do you test streaming?" — Integration tests that read the event stream through the same proxy configuration as production, and assert the first event arrives quickly.