Building AI Features in React

Testing AI UI: Mocked Streams, Contract Tests, Visual States


ShopLens's first tests called the real model. Each test took about 4 seconds, a full run cost about $3, and CI ran about 40 times a day. Worse, the tests were flaky: a snapshot of the summary failed whenever the model chose a different word. The team replaced the model with a mocked fetch that returned a complete JSON body. The tests became fast and stable, and they tested almost nothing that mattered. The split-character bug, the stale-chunk race and the partial-text state had no test at all, because nothing in the tests streamed.

Good tests for AI UI replace the model with something deterministic, but keep the things that make AI UI hard: chunks, timing, malformed replies and many states. This lesson builds those tests for ShopLens.

What each test layer provesReducer rules,pure unit testsHooks withfake byte streamsRoutes witha fake llmEvery statefrom fixturesReal model:evals and smoketopbottomOnly the top layer ever calls the real model.
Fakes must keep what makes AI UI hard — split characters, dropped connections, malformed replies — or the tests prove nothing.

What to test where

LayerWhat it provesToolReal model?
ReducerState rules and stale-event handlingUnit tests (section 3)No
HooksStreaming, decoding, cancel, partial textVitest, Testing Library, fake streamsNo
BFF routesValidation, repair, evidence checks, status codesVitest, supertest, fake llmNo
ComponentsEvery visual state renders correctlyFixtures, snapshots, screenshotsNo
Model qualityPrompts produce good outputThe 200-product eval setYes, on prompt changes and nightly

The rule in the last column is the important one. Unit tests must be fast, free and deterministic, so they never call the model. The model's quality is tested separately, on purpose, with the saved test sets you have seen in the case studies.

A fake stream you control

The browser client reads response.body as a stream of bytes. A fake that produces those bytes, one chunk at a time, can exercise every streaming edge case:

TypeScript
// src/test/fakeStream.tsexport function fakeStreamResponse(  chunks: Array<string | Uint8Array>,  options: { delayMs?: number; failAfter?: number } = {},): Response {  const encoder = new TextEncoder();  let i = 0;  const body = new ReadableStream<Uint8Array>({    async pull(controller) {      if (options.delayMs) await new Promise((r) => setTimeout(r, options.delayMs));      if (i === options.failAfter) return controller.error(new TypeError("network error"));      if (i === chunks.length) return controller.close();      const chunk = chunks[i++];      controller.enqueue(typeof chunk === "string" ? encoder.encode(chunk) : chunk);    },  });  return new Response(body, { status: 200, headers: { "Content-Type": "text/plain; charset=utf-8" } });}

Chunks can be strings or raw bytes, so a test can split a multi-byte character exactly where it wants. failAfter breaks the stream after a given number of chunks, the way a dropped connection does. With Vitest on Node 20 and its jsdom environment, Response, ReadableStream and TextEncoder are all available. Install @testing-library/react together with @testing-library/dom, which recent versions need as a separate peer package.

TSX
// src/hooks/useStreamingSummary.test.tsximport { renderHook, waitFor } from "@testing-library/react";import { afterEach, expect, test, vi } from "vitest";import { fakeStreamResponse } from "../test/fakeStream";import { useStreamingSummary } from "./useStreamingSummary";afterEach(() => vi.unstubAllGlobals());test("decodes a rupee sign split across two chunks", async () => {  const bytes = new TextEncoder().encode("Worth ₹8,999.");  const cut = 7; // "Worth " is 6 bytes and ₹ is the next 3, so this splits ₹  vi.stubGlobal("fetch", vi.fn(async () => fakeStreamResponse([bytes.slice(0, cut), bytes.slice(cut)])));  const { result } = renderHook(() => useStreamingSummary("p-310"));  await waitFor(() => expect(result.current.state.status).toBe("done"));  expect(result.current.state).toMatchObject({ data: "Worth ₹8,999." });});test("keeps partial text when the connection drops", async () => {  vi.stubGlobal("fetch", vi.fn(async () =>    fakeStreamResponse(["Most buyers ", "praise the battery"], { failAfter: 2 })));  const { result } = renderHook(() => useStreamingSummary("p-310"));  await waitFor(() => expect(result.current.state.status).toBe("error"));  expect(result.current.state).toMatchObject({ partial: "Most buyers praise the battery", retryable: true });});

The first test fails if anyone ever decodes chunks separately, which is exactly the bug from section 1. The second proves that a dropped connection becomes a retryable error that keeps its partial text, which is what the footer in section 4 depends on. Both run in milliseconds and give the same result every time.

Contract tests for the BFF

The routes are tested against a fake llm that returns whatever the test queues up. This is where you test malformed replies, which a real model produces too rarely to test on demand.

TypeScript
// server/routes/compare.test.tsimport express from "express";import request from "supertest";import { expect, test, vi } from "vitest";import { compareRoute } from "./compare";const { replies } = vi.hoisted(() => ({ replies: [] as string[] }));vi.mock("../llm", () => ({ llm: { complete: vi.fn(async () => replies.shift() ?? ""), stream: vi.fn() } }));vi.mock("../data", () => ({  getProduct: async (id: string) => ({ id, name: `Product ${id}` }),  getTopReviews: async (productId: string) => [    { id: `r-${productId}`, productId, rating: 5, title: "Good", body: "Battery lasts all week", helpfulVotes: 3 },  ],}));const app = express().use(express.json()).post("/api/compare", compareRoute);const row = { aspect: "Battery", a: "Lasts a week", b: "Two days", better: "a" };test("drops invented evidence and downgrades the verdict", async () => {  replies.push(JSON.stringify({ rows: [{ ...row, evidence: ["r9999"] }] }));  const res = await request(app).post("/api/compare").send({ productIds: ["p-310", "p-422"] });  expect(res.status).toBe(200);  expect(res.body.rows[0]).toMatchObject({ evidence: [], better: "unclear" });});test("repairs once, then answers 502", async () => {  replies.push("Sure! Here is your table:", "still not JSON");  const res = await request(app).post("/api/compare").send({ productIds: ["p-310", "p-501"] });  expect(res.status).toBe(502);});

vi.hoisted creates the replies queue before the mocks are set up, because Vitest moves vi.mock calls to the top of the file. Each test uses a different product pair, since the route caches by pair. The same pattern covers the summary route: a fake stream that yields "NOT_", "ENOUGH_INFO" must produce a 422, and one that yields 100 words must stop at 80. The fake is possible because section 3 hid the provider behind the small Llm interface.

To keep the shared schemas honest, save a few dozen real model replies from the eval runs as JSON fixtures, and add a test that each one still passes its schema. When someone tightens a schema, that test shows at once whether real replies would now be rejected.

Every visual state, on purpose

An AI component has many states, and most of them are rare in manual testing. Nobody sees the "cut short" state unless a connection happens to drop. So ShopLens renders every state from fixtures. To make this easy, the JSX of BuyerSummary moves into a SummaryView component that takes the state as a prop, while BuyerSummary only connects the hook to it.

TSX
// src/components/SummaryView.states.test.tsximport { render } from "@testing-library/react";import { expect, test } from "vitest";import type { RequestState } from "../state/requestMachine";import { SummaryView } from "./SummaryView";const STATES: Record<string, RequestState<string>> = {  idle: { status: "idle" },  loading: { status: "loading", id: 1 },  streaming: { status: "streaming", id: 1, text: "Most buyers praise the battery [r1042]" },  streamingHalfCitation: { status: "streaming", id: 1, text: "Cushions get warm [r09" },  done: { status: "done", id: 1, data: "Most buyers praise the battery [r1042]. Comfort is mixed [r0988]." },  cutShort: { status: "error", id: 1, message: "The stream was interrupted", retryable: true,    partial: "Most buyers praise the battery, which lasts a full work week for many of them." },  fragment: { status: "error", id: 1, message: "The stream was interrupted", retryable: true, partial: "Most buyers" },  notEnough: { status: "error", id: 1, message: "not_enough_reviews", retryable: false, partial: "" },  cancelled: { status: "cancelled", id: 1, partial: "Most buyers praise" },};test.each(Object.entries(STATES))("renders the %s state", (_name, state) => {  const { container } = render(    <SummaryView productId="p-310" state={state} knownIds={new Set(["r1042", "r0988"])}      onStop={() => {}} onRetry={() => {}} />,  );  expect(container).toMatchSnapshot();});

Snapshots catch unintended changes, and a few targeted assertions check the rules that matter most: the Stop button exists only while loading or streaming, a half-received citation never appears as raw text, and the fragment state shows the fallback instead of two words. For layout, the same fixtures feed a states page that a Playwright test screenshots at 360 and 1,280 pixels wide, in light and dark themes. That is where you catch an 80-word summary overflowing its box on a small phone.

Where the real model is tested

The real model is tested for quality, not for code correctness, and not on every commit. Two runs cover it. When a prompt changes, the 200-product eval set runs against the new prompt, and the changes are compared with the old prompt: schema failures, word counts, citation validity, and a human reading of a sample. And every night, a smoke test calls the real routes for three products and checks only that replies are valid and arrive within the latency budget. That catches a provider change or an expired key before shoppers do.

Check your understanding

0 of 3 answered

1.Tests that mock fetch to return a complete JSON body pass, but a split-character bug reaches production. Why?

2.Why do ShopLens unit tests never call the real model?

3.What is the main benefit of rendering every RequestState from fixtures?