Course Content
Building AI Features in React
5 sections · 21 lessons
From One Response to a Stream of Chunks
The ShopLens summary takes about 3 seconds from click to last word. The first token, though, is ready after about 0.8 seconds. Without streaming, the shopper watches a spinner for 3 seconds and then the whole paragraph appears. With streaming, the first words appear at 0.8 seconds and the rest arrive at reading speed.
The total time is the same. What changes is the time to the first useful thing on screen, and that is what people feel. A three-second blank wait feels broken; a paragraph that starts writing itself in under a second feels alive.
This lesson explains what a stream is at the HTTP level, how to read one in the browser, and the traps that make a working stream arrive all at once in production.
How a model produces text
A model generates one token at a time, and each token depends on all the tokens before it. It cannot write the last sentence first. On the larger ShopLens model, output arrives at about 60 tokens per second, so a 110-token summary takes about 1.8 seconds to write after the first token.
Streaming simply forwards each piece as soon as it exists, instead of waiting for the end. The provider streams to your BFF, and your BFF streams to the browser. Nothing about the model changes.
What a chunk is, and what it is not
The browser does not receive tokens. It receives chunks: whatever bytes arrived in one network read. A chunk might be half a word ("bat"), several words, or, in a busy moment, the whole answer. Chunk boundaries depend on network buffers, proxies and timing, and they change from run to run.
Worse, a chunk can end in the middle of a character. In UTF-8, the rupee sign ₹ takes 3 bytes and many emoji take 4. If a chunk ends after the first byte of ₹, decoding that chunk alone produces a broken replacement character. So two rules follow:
- Never parse or interpret a single chunk on its own. Append it to what you have.
- Decode bytes with a decoder that remembers incomplete characters between chunks.
Reading a stream with fetch
In the browser, response.body is a ReadableStream of bytes. Here is the smallest correct loop to read it as text:
1// A minimal reader. Section 3 turns this into streamText() in src/api/client.ts.2export async function readTextStream(3 res: Response,4 onText: (text: string) => void,5): Promise<void> {6 if (!res.body) throw new Error("Response has no body");7 const reader = res.body.getReader();8 const decoder = new TextDecoder(); // UTF-8 by default910 for (;;) {11 const { done, value } = await reader.read();12 if (done) break;13 // stream: true keeps a half-received character for the next chunk14 const text = decoder.decode(value, { stream: true });15 if (text) onText(text);16 }17 const rest = decoder.decode(); // flush anything still buffered18 if (rest) onText(rest);19}The loop asks the reader for the next chunk, which arrives as a Uint8Array of bytes. decoder.decode(value, { stream: true }) turns the bytes into text but holds back any incomplete character until the next call. When the stream ends, one final decode() with no arguments flushes what is left. The caller receives text pieces and appends them to state.
Plain text, SSE or JSON lines
There are three common formats for a stream from your BFF to the browser.
| Format | What travels | Good for | Watch out for |
|---|---|---|---|
| Plain text | Raw text bytes | One stream of prose, like the ShopLens summary | No way to send metadata or a mid-stream error message |
| Server-sent events (SSE) | Lines like event: delta and data: ... | Several event types, reconnection | The built-in EventSource supports only GET requests, so POST needs a hand-written parser |
| JSON lines (NDJSON) | One JSON object per line | Structured events: text, citations, usage | You must split on newlines yourself, since a chunk can end mid-line |
ShopLens uses plain text for the summary because the summary is just prose, and the simplest format is the easiest to test. If you later need to send citations or token counts as separate events, JSON lines is the natural next step. Many teams choose SSE because provider APIs use it; your BFF does not have to copy the provider's format.
Streams need an end, and an exit
A normal response either succeeds or fails. A stream can fail halfway. By the time the model fails, your BFF has already sent a 200 status and half a summary, so it cannot change the status code. In ShopLens, the BFF closes the connection abruptly, and the browser's reader.read() rejects. The UI then shows the partial text with a note, a pattern you will build in section 4.
The user needs an exit too. An AbortController passed to fetch lets the Stop button cancel the stream. When the browser aborts, the connection closes, the BFF notices, and it aborts its own call to the model. That last step matters: output tokens are billed as they are generated, so a shopper who stops at word 20 should not pay for words 21 to 80.
Check your understanding
0 of 3 answered
1.Your stream shows a broken character "�" where a rupee sign should be, but only sometimes. What is the most likely cause?
2.Streaming works locally, but in production the whole summary appears at once after 3 seconds. What should you check first?
3.Why does ShopLens use plain text rather than JSON lines for the summary stream?