Building AI Features in React

What a Model Response Breaks: Shape, Timing, Truth


The first ShopLens prototype asked the model one question per review: "Is this review positive, negative or mixed?" The component then did a switch on the reply to pick a badge colour. In a test on 200 real reviews, the replies included Positive, positive., **Mixed**, and "The review is mostly positive, although the reviewer mentions...". The switch fell through to its default case on 23 of the 200 reviews, about 11 percent.

Nothing crashed. There was no error in the console. The badges were simply missing or wrong, and the only way to find out was to look. This is the typical way AI features fail in a UI: quietly, with a 200 status code.

This lesson names the three things a model reply breaks, shape, timing and truth, plus a fourth that surprises frontend teams: cost per render.

Four silent breaks, each with a 200 status11% of badges missingValidate,then fall back6.9 s at p95Stream,Stop, drop stale"30-hour battery"Sources, honest labels1,000 USD a dayCache per productWhat ShopLens sawWhere the defence livesShapeTimingTruthCostNone of these throws an error; each has to be caught on purpose.
A model call fails quietly, so every break needs a named defence in the frontend or the BFF.

Shape: the reply is a string that might contain what you want

A model returns text. Even when you ask for one word or for JSON, you get a string that usually contains what you asked for. Here is the naive prototype:

TSX
// The prototype. Do not ship this.function badgeFor(reply: string) {  switch (reply) {    case "positive": return <Badge tone="green">Positive</Badge>;    case "negative": return <Badge tone="red">Negative</Badge>;    case "mixed":    return <Badge tone="amber">Mixed</Badge>;    default:         return null; // 11% of reviews ended up here  }}

These are the shape failures you will see, roughly from most to least common:

  • Extra words: "Positive." or "The sentiment is positive".
  • Formatting: Markdown bold, or JSON wrapped in a code fence with the word json.
  • Different case or synonyms: "Positive", "favourable", "neutral".
  • Extra or missing keys in JSON, or a number sent as a string: "rating": "4".
  • Truncation: the reply hits the token limit and the JSON stops halfway, with no closing brace.
  • A new value you never listed, such as a fourth sentiment called "neutral".

Better prompts reduce these (section 2), and provider features that constrain output to a schema reduce them further. But none of them make the check unnecessary, because truncation and new values can still happen. The rule is simple: the string is untrusted until your code has parsed and validated it.

Timing: seconds, and never the same twice

A model writes one token at a time. Total time is roughly the time to the first token plus the number of output tokens divided by the speed. For the ShopLens summary on the larger model, a typical run looks like this:

StageTypical time
Network and queueing at the provider100–300 ms
Reading the 3,600-token prompt, until the first token400–900 ms
Writing 110 tokens at about 60 tokens per secondabout 1.8 s
Totalabout 2.5–3 s

And it varies. Across one day of ShopLens traffic, the summary took 2.4 seconds at the median and 6.9 seconds at the 95th percentile. Busy periods at the provider, longer answers and retries after a rate-limit error all stretch the tail. Occasionally a call fails after 30 seconds.

For the UI, this has three consequences. A spinner alone is not enough for 3 seconds, and far from enough for 7. The user needs a way to cancel. And because calls overlap, you must ignore responses that arrive for a product the user has already left.

Truth: fluent and wrong

In an early test, the ShopLens summary said: "Buyers report that the battery lasts about 30 hours." The product specification says 24 hours. The reviews that mention battery life say 18 to 22 hours. The number 30 appeared nowhere.

This is called a hallucination: the model produced a plausible statement that the input does not support. It happens because a model is trained to produce likely text, not to check facts. "About 30 hours" is a very likely phrase in headphone reviews in general. The model has no built-in sense that it must come from these 40 reviews.

For a frontend engineer, the lesson is about presentation. Your UI gives every sentence it shows the authority of your brand. So a model-written sentence needs different treatment from a database value:

  • Label it as a summary of buyers' opinions ("What buyers say"), not as a fact.
  • Link claims to their sources, so a shopper can check them (section 4).
  • Avoid showing numbers that you did not compute yourself.
  • Give the shopper a way to report a wrong summary.

A fourth break: cost per render

Frontend engineers are used to calls that are free at the margin. Re-fetching the reviews list costs a database query. A model call costs money every time, and React makes it easy to call things more often than you meant to.

Put numbers on it. Kestrel X2 and the rest of the store get 50,000 product page views a day. If every view generates a fresh $0.02 summary, that is $1,000 a day. If summaries are cached per product for 24 hours, and the store has 3,000 products, the worst case is 3,000 × $0.02 = $60 a day. Same feature, same model, a seventeen-times difference, decided entirely by frontend and BFF design. Common frontend causes of wasted calls include an effect with an unstable dependency that runs on every render, a request fired on every keystroke, and development mode in React 18, where Strict Mode runs effects twice.

BreakWhat your UI needsWhere in this course
ShapeParse and validate at the boundary; a fallback when invalidSection 2, lesson 4
TimingStreaming, skeletons, a Stop button, stale-response guardsSections 3 and 4
TruthHonest labels, sources, feedbackSection 4
CostServer-side cache, dedupe, debounce, limitsSection 5

Check your understanding

0 of 3 answered

1.A reply for the sentiment badge is "Positive." and your code expects "positive". What is the best response?

2.The summary takes 2.4 s at the median and 6.9 s at the 95th percentile. What does that mean for the UI?

3.Why did the summary say "about 30 hours" when no review mentioned it?