Building AI Features in React

Showing Confidence and Sources Honestly


A product manager asked for a small percentage next to each sentiment badge: "92% confident". It seemed easy. The team added a confidence field from 0 to 100 to the badge prompt and ran it on the 200 test reviews. Of the answers, 87 percent were 85, 90 or 95. And badges marked 95 were wrong about as often as badges marked 85: roughly 8 percent of the time for both.

The numbers looked precise and meant nothing. Showing them would have been worse than showing nothing, because shoppers trust a percentage. This lesson shows what to show instead: signals you can actually compute, and sources the shopper can check.

Where confidence should come fromA number the model wrote• Clusters on 85, 90 and 95• Wrong as often at 95 as at 85• Looks precise, predicts nothingSignals your code computes• Label agrees with the star rating• Citations point at real reviews• Scope: 40 of 214 reviews
Hiding badges that disagree with the stars dropped 9 percent of badges and about 60 percent of the wrong ones.

A stated confidence is just more generated text

When you ask a model for a confidence score, it writes a number the same way it writes any other word: as a likely continuation. It is not reading out an internal probability. So it clusters on familiar round numbers, and it is usually high, because confident-sounding text is common in training data.

Some APIs expose token probabilities, and for single-token classifications these can carry real signal. They are not available from every provider or model, and turning them into a calibrated confidence needs a labelled test set. For most UI work, the practical rule is simple: do not display a confidence number that the model wrote about itself.

Confidence you can compute

Better signals come from checks your code can run against data it trusts.

FeatureSignalHow it is computed
Sentiment badgeLabel agrees with the star ratingA 5-star review labelled negative is suspicious
SummaryHow many reviews stand behind it"Based on 40 of 214 reviews", counted by the server
SummaryClaims with a verified citationCitation ids that exist in the reviews sent
Compare tableEvidence per rowRows with no real evidence become "unclear" (section 3)

The badge check is a few lines:

TypeScript
// src/lib/confidence.tsimport type { Review, SentimentLabel } from "../../shared/schemas";export function agreesWithStars(label: SentimentLabel, rating: Review["rating"]): boolean {  if (label === "positive") return rating >= 4;  if (label === "negative") return rating <= 2;  return rating >= 2 && rating <= 4; // mixed}

On the test set, the badge was wrong on 6 percent of reviews overall, but on 41 percent of the reviews where it disagreed with the stars. Disagreements were 9 percent of reviews. So ShopLens simply shows no badge when the label and the stars disagree. It loses 9 percent of badges and removes about 60 percent of the wrong ones. The stars remain on every review, so the shopper loses nothing they could not already see.

This is the general pattern: find a cheap, trusted signal that correlates with the model being wrong, and use it to hide or soften the output, not to decorate it with a number.

Sources the shopper can check

The summary prompt already asks the model to cite review ids after each claim, like [r1042]. A citation is only useful if the shopper can follow it, and only honest if it points at a real review. The CitedText component turns markers into small numbered buttons, drops any id it does not know, and hides a half-received marker at the end of a streaming text.

TSX
// src/components/CitedText.tsximport { Fragment } from "react";const CITATION = /\[(r\d+)\]/;       // one capture group: split() keeps the idconst PARTIAL_TAIL = /\[r?\d*$/;      // "[", "[r" or "[r10" at the very endtype Props = {  text: string;  streaming: boolean;  knownIds: ReadonlySet<string>;  onCite: (reviewId: string) => void;};export function CitedText({ text, streaming, knownIds, onCite }: Props) {  const clean = streaming ? text.replace(PARTIAL_TAIL, "") : text;  const parts = clean.split(CITATION); // even indexes: text, odd indexes: ids  const numbers = new Map<string, number>();  return (    <p>      {parts.map((part, i) => {        if (i % 2 === 0) return <Fragment key={i}>{part}</Fragment>;        if (!knownIds.has(part)) return null; // an invented id is never shown        const n = numbers.get(part) ?? numbers.size + 1;        numbers.set(part, n);        return (          <button key={i} type="button" className="cite" aria-label={`Source: review ${n}`}            onClick={() => onCite(part)}>            {n}          </button>        );      })}    </p>  );}

String.split with a capturing group puts the captured ids at the odd indexes of the result, so the loop alternates between plain text and ids. The same review cited twice gets the same number. Clicking a number opens that review in a dialog, with the quoted sentence highlighted if the shopper wants to check the claim.

Where do the known ids come from? ShopLens's review list already loads the product's 40 most helpful reviews, the same query the summary route uses, so the panel passes those ids in. The alternative is to filter citations on the server as the stream passes through, which is stricter but needs the same hold-and-release logic you wrote for the sentinel. Either way, an id that does not exist never reaches the screen.

In BuyerSummary, the plain paragraph from the previous lesson becomes <CitedText text={text} streaming={state.status === "streaming"} knownIds={ids} onCite={openReview} />. Everything is still rendered as React text and elements, never as HTML.

Words that do not over-claim

Wording is part of honesty. The same summary can be presented in a way that invites checking or in a way that demands trust.

Over-claimsHonest
"The truth about Kestrel X2""What buyers say"
"Battery: 30 hours""Many buyers report a full work week per charge"
A summary with no label"Summary of 40 of 214 reviews, written by AI"
"Best for comfort" in the table"Buyers rate Wren Air better for comfort"

The line "Summary of 40 of 214 reviews, written by AI" does two jobs. It tells the shopper the scope, which the server computed, not the model. And it discloses that the text was generated, which shoppers increasingly expect and which rules in several countries are moving towards requiring.

Check your understanding

0 of 3 answered

1.The model returns "confidence": 95 for a badge. What should the UI do with it?

2.While streaming, the summary text currently ends with "...warm after two hours [r09". What does CitedText render?

3.Why does ShopLens show "Summary of 40 of 214 reviews" under the summary?