Course Content
Building AI Features in React
5 sections · 21 lessons
Testing AI UI: Mocked Streams, Contract Tests, Visual States
ShopLens's first tests called the real model. Each test took about 4 seconds, a full run cost about $3, and CI ran about 40 times a day. Worse, the tests were flaky: a snapshot of the summary failed whenever the model chose a different word. The team replaced the model with a mocked fetch that returned a complete JSON body. The tests became fast and stable, and they tested almost nothing that mattered. The split-character bug, the stale-chunk race and the partial-text state had no test at all, because nothing in the tests streamed.
Good tests for AI UI replace the model with something deterministic, but keep the things that make AI UI hard: chunks, timing, malformed replies and many states. This lesson builds those tests for ShopLens.
What to test where
| Layer | What it proves | Tool | Real model? |
|---|---|---|---|
| Reducer | State rules and stale-event handling | Unit tests (section 3) | No |
| Hooks | Streaming, decoding, cancel, partial text | Vitest, Testing Library, fake streams | No |
| BFF routes | Validation, repair, evidence checks, status codes | Vitest, supertest, fake llm | No |
| Components | Every visual state renders correctly | Fixtures, snapshots, screenshots | No |
| Model quality | Prompts produce good output | The 200-product eval set | Yes, on prompt changes and nightly |
The rule in the last column is the important one. Unit tests must be fast, free and deterministic, so they never call the model. The model's quality is tested separately, on purpose, with the saved test sets you have seen in the case studies.
A fake stream you control
The browser client reads response.body as a stream of bytes. A fake that produces those bytes, one chunk at a time, can exercise every streaming edge case:
1// src/test/fakeStream.ts2export function fakeStreamResponse(3 chunks: Array<string | Uint8Array>,4 options: { delayMs?: number; failAfter?: number } = {},5): Response {6 const encoder = new TextEncoder();7 let i = 0;8 const body = new ReadableStream<Uint8Array>({9 async pull(controller) {10 if (options.delayMs) await new Promise((r) => setTimeout(r, options.delayMs));11 if (i === options.failAfter) return controller.error(new TypeError("network error"));12 if (i === chunks.length) return controller.close();13 const chunk = chunks[i++];14 controller.enqueue(typeof chunk === "string" ? encoder.encode(chunk) : chunk);15 },16 });17 return new Response(body, { status: 200, headers: { "Content-Type": "text/plain; charset=utf-8" } });18}Chunks can be strings or raw bytes, so a test can split a multi-byte character exactly where it wants. failAfter breaks the stream after a given number of chunks, the way a dropped connection does. With Vitest on Node 20 and its jsdom environment, Response, ReadableStream and TextEncoder are all available. Install @testing-library/react together with @testing-library/dom, which recent versions need as a separate peer package.
1// src/hooks/useStreamingSummary.test.tsx2import { renderHook, waitFor } from "@testing-library/react";3import { afterEach, expect, test, vi } from "vitest";4import { fakeStreamResponse } from "../test/fakeStream";5import { useStreamingSummary } from "./useStreamingSummary";67afterEach(() => vi.unstubAllGlobals());89test("decodes a rupee sign split across two chunks", async () => {10 const bytes = new TextEncoder().encode("Worth ₹8,999.");11 const cut = 7; // "Worth " is 6 bytes and ₹ is the next 3, so this splits ₹12 vi.stubGlobal("fetch", vi.fn(async () => fakeStreamResponse([bytes.slice(0, cut), bytes.slice(cut)])));13 const { result } = renderHook(() => useStreamingSummary("p-310"));14 await waitFor(() => expect(result.current.state.status).toBe("done"));15 expect(result.current.state).toMatchObject({ data: "Worth ₹8,999." });16});1718test("keeps partial text when the connection drops", async () => {19 vi.stubGlobal("fetch", vi.fn(async () =>20 fakeStreamResponse(["Most buyers ", "praise the battery"], { failAfter: 2 })));21 const { result } = renderHook(() => useStreamingSummary("p-310"));22 await waitFor(() => expect(result.current.state.status).toBe("error"));23 expect(result.current.state).toMatchObject({ partial: "Most buyers praise the battery", retryable: true });24});The first test fails if anyone ever decodes chunks separately, which is exactly the bug from section 1. The second proves that a dropped connection becomes a retryable error that keeps its partial text, which is what the footer in section 4 depends on. Both run in milliseconds and give the same result every time.
Contract tests for the BFF
The routes are tested against a fake llm that returns whatever the test queues up. This is where you test malformed replies, which a real model produces too rarely to test on demand.
1// server/routes/compare.test.ts2import express from "express";3import request from "supertest";4import { expect, test, vi } from "vitest";5import { compareRoute } from "./compare";67const { replies } = vi.hoisted(() => ({ replies: [] as string[] }));8vi.mock("../llm", () => ({ llm: { complete: vi.fn(async () => replies.shift() ?? ""), stream: vi.fn() } }));9vi.mock("../data", () => ({10 getProduct: async (id: string) => ({ id, name: `Product ${id}` }),11 getTopReviews: async (productId: string) => [12 { id: `r-${productId}`, productId, rating: 5, title: "Good", body: "Battery lasts all week", helpfulVotes: 3 },13 ],14}));1516const app = express().use(express.json()).post("/api/compare", compareRoute);17const row = { aspect: "Battery", a: "Lasts a week", b: "Two days", better: "a" };1819test("drops invented evidence and downgrades the verdict", async () => {20 replies.push(JSON.stringify({ rows: [{ ...row, evidence: ["r9999"] }] }));21 const res = await request(app).post("/api/compare").send({ productIds: ["p-310", "p-422"] });22 expect(res.status).toBe(200);23 expect(res.body.rows[0]).toMatchObject({ evidence: [], better: "unclear" });24});2526test("repairs once, then answers 502", async () => {27 replies.push("Sure! Here is your table:", "still not JSON");28 const res = await request(app).post("/api/compare").send({ productIds: ["p-310", "p-501"] });29 expect(res.status).toBe(502);30});vi.hoisted creates the replies queue before the mocks are set up, because Vitest moves vi.mock calls to the top of the file. Each test uses a different product pair, since the route caches by pair. The same pattern covers the summary route: a fake stream that yields "NOT_", "ENOUGH_INFO" must produce a 422, and one that yields 100 words must stop at 80. The fake is possible because section 3 hid the provider behind the small Llm interface.
To keep the shared schemas honest, save a few dozen real model replies from the eval runs as JSON fixtures, and add a test that each one still passes its schema. When someone tightens a schema, that test shows at once whether real replies would now be rejected.
Every visual state, on purpose
An AI component has many states, and most of them are rare in manual testing. Nobody sees the "cut short" state unless a connection happens to drop. So ShopLens renders every state from fixtures. To make this easy, the JSX of BuyerSummary moves into a SummaryView component that takes the state as a prop, while BuyerSummary only connects the hook to it.
1// src/components/SummaryView.states.test.tsx2import { render } from "@testing-library/react";3import { expect, test } from "vitest";4import type { RequestState } from "../state/requestMachine";5import { SummaryView } from "./SummaryView";67const STATES: Record<string, RequestState<string>> = {8 idle: { status: "idle" },9 loading: { status: "loading", id: 1 },10 streaming: { status: "streaming", id: 1, text: "Most buyers praise the battery [r1042]" },11 streamingHalfCitation: { status: "streaming", id: 1, text: "Cushions get warm [r09" },12 done: { status: "done", id: 1, data: "Most buyers praise the battery [r1042]. Comfort is mixed [r0988]." },13 cutShort: { status: "error", id: 1, message: "The stream was interrupted", retryable: true,14 partial: "Most buyers praise the battery, which lasts a full work week for many of them." },15 fragment: { status: "error", id: 1, message: "The stream was interrupted", retryable: true, partial: "Most buyers" },16 notEnough: { status: "error", id: 1, message: "not_enough_reviews", retryable: false, partial: "" },17 cancelled: { status: "cancelled", id: 1, partial: "Most buyers praise" },18};1920test.each(Object.entries(STATES))("renders the %s state", (_name, state) => {21 const { container } = render(22 <SummaryView productId="p-310" state={state} knownIds={new Set(["r1042", "r0988"])}23 onStop={() => {}} onRetry={() => {}} />,24 );25 expect(container).toMatchSnapshot();26});Snapshots catch unintended changes, and a few targeted assertions check the rules that matter most: the Stop button exists only while loading or streaming, a half-received citation never appears as raw text, and the fragment state shows the fallback instead of two words. For layout, the same fixtures feed a states page that a Playwright test screenshots at 360 and 1,280 pixels wide, in light and dark themes. That is where you catch an 80-word summary overflowing its box on a small phone.
Where the real model is tested
The real model is tested for quality, not for code correctness, and not on every commit. Two runs cover it. When a prompt changes, the 200-product eval set runs against the new prompt, and the changes are compared with the old prompt: schema failures, word counts, citation validity, and a human reading of a sample. And every night, a smoke test calls the real routes for three products and checks only that replies are valid and arrive within the latency budget. That catches a provider change or an expired key before shoppers do.
Check your understanding
0 of 3 answered
1.Tests that mock fetch to return a complete JSON body pass, but a split-character bug reaches production. Why?
2.Why do ShopLens unit tests never call the real model?
3.What is the main benefit of rendering every RequestState from fixtures?