Building AI Features in React

Loading and Streaming That Feel Fast


After section 3, the ShopLens summary showed its first word in 0.8 seconds. In a moderated test with eight shoppers, most still described it as "a bit jumpy". The recordings showed why. The panel changed height three times: when the skeleton appeared, as the text grew past the skeleton's height, and when the buttons appeared at the end. The text arrived in bursts, nothing for 300 ms and then a dozen words at once, so it stuttered instead of flowing. And one participant using a screen reader heard fragments of sentences read out as they arrived.

None of these are model problems. The model was as fast as it will get. This lesson is about the frontend work that decides whether the same 3 seconds feel smooth or broken.

Match the waiting pattern to the waitShow nothing newcachedsummary, 30 msSkeletonafter 150 msbefore thefirst wordStream, saywhat runscompare table, 5 sLet them continuenot offeredPatternShopLens caseUnder 300 ms0.3 to 1 s1 to 10 sOver 10 s
The same three seconds feel smooth or broken depending on reserved space, a delayed skeleton and one announcement at the end.

Match the pattern to the wait

Different waits need different treatment. These thresholds are rough, but they match what most teams find in testing.

Expected waitPatternShopLens example
Under 300 msShow nothing new; a flash of a loader is worse than a short pauseA cached summary returned in 30 ms
0.3 to 1 sA skeleton shaped like the resultThe summary before its first word
1 to 10 sA skeleton plus streaming, or a line saying what is happeningThe 5-second compare table: "Reading 50 reviews of both products…"
Over 10 sLet the user continue; notify when doneA full report on all 214 reviews, which ShopLens does not offer

The first row catches teams out. When a cached summary arrives in 30 ms, a skeleton that appears for one frame and vanishes looks like a glitch. The ShopLens skeleton fades in with a CSS animation delay of 150 ms. If the answer arrives first, the skeleton is never seen. This costs no JavaScript at all.

Reserve the space

The summary box has a minimum height of four lines of text, which is 96 pixels at a 24-pixel line height. The skeleton has four grey bars of that line height. The streaming text fills the same box. The Stop button lives in the header row, next to the title, where it is present from the first moment and never moves. The reviews below the panel do not move until the summary grows past four lines, which with a 60-word target is rare.

This matters beyond comfort. Layout shift is measured by Cumulative Layout Shift, one of the Core Web Vitals, and a panel that pushes the Add to Cart button down while the shopper reaches for it causes mis-taps. Before these changes, the product page's layout shift score on phones was 0.18, which counts as "needs improvement". After reserving space, it was 0.01.

Smooth the stream

Chunks arrive in bursts because of how providers batch tokens and how networks deliver packets. Rendering each chunk exactly as it arrives is honest but jerky. A common fix is to reveal the text a little at a time on each animation frame, catching up faster when far behind so that the display never lags the real stream by more than a fraction of a second.

TypeScript
// src/hooks/useSmoothedText.tsimport { useEffect, useRef, useState } from "react";/** Reveals `target` a little per frame, so bursty chunks read as a steady flow. */export function useSmoothedText(target: string, active: boolean): string {  const [shown, setShown] = useState(target);  const shownRef = useRef(target);  useEffect(() => {    // Not streaming, or a brand-new text: show everything immediately.    if (!active || !target.startsWith(shownRef.current)) {      shownRef.current = target;      setShown(target);      return;    }    let frame = 0;    const tick = () => {      const backlog = target.length - shownRef.current.length;      if (backlog <= 0) return;      const step = Math.max(1, Math.ceil(backlog / 8)); // reveal faster when far behind      shownRef.current = target.slice(0, shownRef.current.length + step);      setShown(shownRef.current);      frame = requestAnimationFrame(tick);    };    frame = requestAnimationFrame(tick);    return () => cancelAnimationFrame(frame);  }, [target, active]);  return shown;}

Each frame reveals one eighth of the text that is waiting, and at least one character. A burst of 40 characters appears over about 150 ms instead of in a single frame, and a large backlog is cleared quickly because the step grows with it. When streaming ends, or when a new request starts with different text, the hook shows the full text at once. That means a cached summary, which arrives complete, is shown instantly instead of being "typed" for effect.

The trade-off is honesty against smoothness. Every millisecond of smoothing is a millisecond the shopper sees less than you have. Keep the lag small, never invent a typing effect for text that is already complete, and for users who set "reduce motion" in their system, pass active as false so the text appears as it arrives.

Announce once, not every chunk

A live region is an element marked with aria-live, and screen readers read out its changes. Putting it on the streaming paragraph makes the screen reader announce every chunk, dozens of times per summary, often cutting itself off mid-word. The better pattern is a separate, visually hidden live region that stays empty while text streams and says one short sentence when the summary is complete. The summary box itself gets aria-busy while loading or streaming, which tells assistive technology that the content is still changing.

Here is the summary component with all of this in place, built on the state machine from section 3:

TSX
// src/components/BuyerSummary.tsximport { useId } from "react";import { useSmoothedText } from "../hooks/useSmoothedText";import { useStreamingSummary } from "../hooks/useStreamingSummary";import { visibleText } from "../state/requestMachine";import { SummaryFooter } from "./SummaryFooter";import { SummarySkeleton } from "./SummarySkeleton";export function BuyerSummary({ productId }: { productId: string }) {  const { state, start, stop } = useStreamingSummary(productId);  const titleId = useId();  const busy = state.status === "loading" || state.status === "streaming";  const text = useSmoothedText(visibleText(state), state.status === "streaming");  return (    <section className="buyer-summary" aria-labelledby={titleId} aria-busy={busy}>      <div className="buyer-summary__head">        <h2 id={titleId}>What buyers say</h2>        {busy && <button type="button" onClick={stop}>Stop</button>}      </div>      <div className="buyer-summary__body">{/* min-height: four lines */}        {state.status === "loading" ? (          <SummarySkeleton lines={4} />        ) : (          <p>{text}{state.status === "streaming" && <span className="caret" aria-hidden="true" />}</p>        )}      </div>      <p className="visually-hidden" aria-live="polite">        {state.status === "done" ? "Buyer summary ready." : ""}      </p>      <SummaryFooter productId={productId} state={state} onRetry={start} />    </section>  );}

SummaryFooter is where the controls for finished, failed and stopped summaries live; the next three lessons fill it in. Notice that the component contains no timing logic of its own. The hook owns the request, the reducer owns the rules, and useSmoothedText owns presentation timing, so each can be changed and tested alone.

Start earlier, when it is cheap to

The fastest wait is one that has already happened. ShopLens starts the summary request when the panel comes within 300 pixels of the viewport, using an IntersectionObserver, instead of when the shopper reaches it. On a phone, that hides most of the first-word delay behind scrolling. The cost is paying for summaries some shoppers never scroll to, which is only acceptable because summaries are cached per product (section 5). Without a cache, starting early would multiply your bill.

Check your understanding

0 of 3 answered

1.A cached summary returns in 30 ms, but shoppers see a skeleton flash for one frame. What is the simplest fix?

2.Why is aria-live not placed on the streaming paragraph?

3.What does useSmoothedText do when streaming ends?