Course Content
Building AI Features in React
5 sections · 21 lessons
Showing Confidence and Sources Honestly
A product manager asked for a small percentage next to each sentiment badge: "92% confident". It seemed easy. The team added a confidence field from 0 to 100 to the badge prompt and ran it on the 200 test reviews. Of the answers, 87 percent were 85, 90 or 95. And badges marked 95 were wrong about as often as badges marked 85: roughly 8 percent of the time for both.
The numbers looked precise and meant nothing. Showing them would have been worse than showing nothing, because shoppers trust a percentage. This lesson shows what to show instead: signals you can actually compute, and sources the shopper can check.
A stated confidence is just more generated text
When you ask a model for a confidence score, it writes a number the same way it writes any other word: as a likely continuation. It is not reading out an internal probability. So it clusters on familiar round numbers, and it is usually high, because confident-sounding text is common in training data.
Some APIs expose token probabilities, and for single-token classifications these can carry real signal. They are not available from every provider or model, and turning them into a calibrated confidence needs a labelled test set. For most UI work, the practical rule is simple: do not display a confidence number that the model wrote about itself.
Confidence you can compute
Better signals come from checks your code can run against data it trusts.
| Feature | Signal | How it is computed |
|---|---|---|
| Sentiment badge | Label agrees with the star rating | A 5-star review labelled negative is suspicious |
| Summary | How many reviews stand behind it | "Based on 40 of 214 reviews", counted by the server |
| Summary | Claims with a verified citation | Citation ids that exist in the reviews sent |
| Compare table | Evidence per row | Rows with no real evidence become "unclear" (section 3) |
The badge check is a few lines:
1// src/lib/confidence.ts2import type { Review, SentimentLabel } from "../../shared/schemas";34export function agreesWithStars(label: SentimentLabel, rating: Review["rating"]): boolean {5 if (label === "positive") return rating >= 4;6 if (label === "negative") return rating <= 2;7 return rating >= 2 && rating <= 4; // mixed8}On the test set, the badge was wrong on 6 percent of reviews overall, but on 41 percent of the reviews where it disagreed with the stars. Disagreements were 9 percent of reviews. So ShopLens simply shows no badge when the label and the stars disagree. It loses 9 percent of badges and removes about 60 percent of the wrong ones. The stars remain on every review, so the shopper loses nothing they could not already see.
This is the general pattern: find a cheap, trusted signal that correlates with the model being wrong, and use it to hide or soften the output, not to decorate it with a number.
Sources the shopper can check
The summary prompt already asks the model to cite review ids after each claim, like [r1042]. A citation is only useful if the shopper can follow it, and only honest if it points at a real review. The CitedText component turns markers into small numbered buttons, drops any id it does not know, and hides a half-received marker at the end of a streaming text.
1// src/components/CitedText.tsx2import { Fragment } from "react";34const CITATION = /\[(r\d+)\]/; // one capture group: split() keeps the id5const PARTIAL_TAIL = /\[r?\d*$/; // "[", "[r" or "[r10" at the very end67type Props = {8 text: string;9 streaming: boolean;10 knownIds: ReadonlySet<string>;11 onCite: (reviewId: string) => void;12};1314export function CitedText({ text, streaming, knownIds, onCite }: Props) {15 const clean = streaming ? text.replace(PARTIAL_TAIL, "") : text;16 const parts = clean.split(CITATION); // even indexes: text, odd indexes: ids17 const numbers = new Map<string, number>();1819 return (20 <p>21 {parts.map((part, i) => {22 if (i % 2 === 0) return <Fragment key={i}>{part}</Fragment>;23 if (!knownIds.has(part)) return null; // an invented id is never shown24 const n = numbers.get(part) ?? numbers.size + 1;25 numbers.set(part, n);26 return (27 <button key={i} type="button" className="cite" aria-label={`Source: review ${n}`}28 onClick={() => onCite(part)}>29 {n}30 </button>31 );32 })}33 </p>34 );35}String.split with a capturing group puts the captured ids at the odd indexes of the result, so the loop alternates between plain text and ids. The same review cited twice gets the same number. Clicking a number opens that review in a dialog, with the quoted sentence highlighted if the shopper wants to check the claim.
Where do the known ids come from? ShopLens's review list already loads the product's 40 most helpful reviews, the same query the summary route uses, so the panel passes those ids in. The alternative is to filter citations on the server as the stream passes through, which is stricter but needs the same hold-and-release logic you wrote for the sentinel. Either way, an id that does not exist never reaches the screen.
In BuyerSummary, the plain paragraph from the previous lesson becomes <CitedText text={text} streaming={state.status === "streaming"} knownIds={ids} onCite={openReview} />. Everything is still rendered as React text and elements, never as HTML.
Words that do not over-claim
Wording is part of honesty. The same summary can be presented in a way that invites checking or in a way that demands trust.
| Over-claims | Honest |
|---|---|
| "The truth about Kestrel X2" | "What buyers say" |
| "Battery: 30 hours" | "Many buyers report a full work week per charge" |
| A summary with no label | "Summary of 40 of 214 reviews, written by AI" |
| "Best for comfort" in the table | "Buyers rate Wren Air better for comfort" |
The line "Summary of 40 of 214 reviews, written by AI" does two jobs. It tells the shopper the scope, which the server computed, not the model. And it discloses that the text was generated, which shoppers increasingly expect and which rules in several countries are moving towards requiring.
Check your understanding
0 of 3 answered
1.The model returns "confidence": 95 for a badge. What should the UI do with it?
2.While streaming, the summary text currently ends with "...warm after two hours [r09". What does CitedText render?
3.Why does ShopLens show "Summary of 40 of 214 reviews" under the summary?