Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 8: Streaming Response Integration
Scenario: a LangChain RAG app needs to stream answers to a web client, and the retrieval step sits in front of generation. How do you build it?
What you need to know
chain.stream() yields the output of the last step. In a RAG chain, that means you see nothing during retrieval, then tokens. astream_events yields events from every step, each labelled with its type, so the client can show progress: "found 4 sources", then the answer as it is written.
The event loop
1async def events(question: str):2 async for ev in chain.astream_events({"question": question}, version="v2"):3 kind = ev["event"]4 if kind == "on_retriever_end":5 docs = ev["data"]["output"]6 yield sse("sources", [d.metadata.get("source") for d in docs])7 elif kind == "on_chat_model_stream":8 yield sse("token", ev["data"]["chunk"].content)9 yield sse("done", {})Sources arrive the moment retrieval ends, often a second before the first token, so the user sees progress immediately. For LangGraph agents, graph.astream(..., stream_mode="messages") streams model tokens with the node that produced them, which is the equivalent pattern.
The FastAPI side
1from fastapi.responses import StreamingResponse23@app.get("/ask")4async def ask(q: str):5 return StreamingResponse(events(q), media_type="text/event-stream",6 headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})X-Accel-Buffering: no tells nginx not to buffer the response. Without it, or the equivalent setting on your load balancer, the proxy collects the whole stream and releases it at the end. Everything works on a laptop and "doesn't stream" in production.
Failure modes to handle explicitly
| Failure | What to do |
|---|---|
| Client disconnects mid-stream | Catch the cancellation, save the partial answer and its cost, stop generation |
| Error after tokens were sent | Send an error event; the HTTP status is already 200 and cannot change |
| Moderation needed | Check each sentence before release, not the final buffer |
| Slow first token | Send a "searching" event at once so the UI isn't blank |
SSE or WebSockets?
Server-sent events
- One-way, server to client
- Plain HTTP; works through most proxies
- Automatic reconnect in browsers
- Right for chat answers
WebSockets
- Two-way, both directions
- Needs upgrade support at every hop
- More connection state to manage
- Right for live collaboration or voice
A real-life example
Scenario (illustrative numbers). A university's admissions assistant shows a spinner for 5 to 6 seconds, then the full answer. Engineers added streaming, which works locally, but in production the answer still appears all at once.
The cause is nginx buffering the response. Adding X-Accel-Buffering: no fixes it. Switching from stream to astream_events lets them send sources after about 400 ms and the first token after about 1.1 s. They also save partial answers on disconnect: about 7% of users close the tab mid-answer, and those conversations are now resumable. In a survey, "the assistant feels fast" rises from 41% to 78%, though total generation time barely changed.
Follow-up questions to expect
- "Why not WebSockets?" — SSE is simpler and enough for one-way streaming; choose WebSockets when the client must send data during the stream.
- "How do you stream agent tool progress?" — Emit events for tool start and end ("checking your order...") from
astream_eventsor LangGraph's stream modes. - "How do you test streaming?" — Integration tests that read the event stream through the same proxy configuration as production, and assert the first event arrives quickly.