- MantraMindAI
- Blog
- AI Applications & Strategy
AI product management when correctness is a distribution
Jai Rao
August 22, 202618 min read
AI features break the deterministic product playbook. How to scope, design the failure path, own the eval set, and measure a feature that is only right most of the time.
Shipping software used to come with a comforting property: you could tell whether it worked. The export button either produced a CSV or it didn't. A bug was a deviation from a written spec, and the spec existed before anyone wrote code. Acceptance criteria were binary, QA signed off, and "works" was a state the product could be in and stay in.
An AI feature does not have that property. Feed the same support thread to the same summariser twice and you may get two different summaries, both defensible. Feed it a thread written half in English and half in Hindi by someone who is furious, and you may get something confidently wrong. The feature isn't broken — it is doing exactly what a probabilistic system does. "Does it work?" has stopped being a yes-or-no question and become a distribution: mostly good, sometimes mediocre, occasionally wrong, and once in a while wrong in a way that makes you wince. Nearly every hard decision in AI product management is downstream of that single fact, so it pays to be precise about what it changes.
The spec you can no longer write
A deterministic spec describes a mapping: for this input, produce this output. You can enumerate the cases because the cases are finite and you chose them. For a generative or predictive feature the input space is whatever users type, and the output space is unbounded. You cannot enumerate anything. What you can specify is the shape of the behaviour and what the product does in each region of that shape.
In practice a usable spec for an AI feature has four parts. A quality floor, stated as an observable user action on a representative sample of inputs rather than as a rubric score — "the agent sends the draft after changing fewer than a quarter of it" is testable, "the summary is high quality" is not. A route for low confidence, because some fraction of inputs will be outside what you evaluated. A route for failure, because the provider will time out and the schema will occasionally come back malformed. And a bound on the worst case: the most damaging thing a wrong output can cause, and what stands between the model and that damage.
The knock-on effect is that QA changes character. You cannot test your way to certainty on a distribution; you can only sample it. Regression testing becomes running a fixed set of inputs and comparing aggregate behaviour to last week's, which means somebody has to own that set of inputs. More on that below, because it turns out to be a product job rather than an engineering one.
Where 85% is a gift and where it is a liability
Two questions decide whether a workflow is a good home for a probabilistic feature, and they are worth asking before any technical feasibility work.
First: is an answer that is right most of the time worth more than no answer? For drafting, triage, extraction, ranking, and search, yes — obviously yes. A first draft that needs editing still beats a blank page, and a queue sorted imperfectly still beats a queue sorted by arrival time. For anything where the user must be able to rely on the output without checking it, no. If the output can't be trusted, the user has to verify it; if they have to verify all of it, verification costs roughly what doing the work cost, and you have shipped a feature that moves effort around instead of removing it.
Second: when it is wrong, what does the error cost and who notices? The cheap-and-visible case looks like this: the output appears in front of a person who can judge it in seconds, nothing has happened yet, and fixing it is one edit. The expensive-and-invisible case looks like this: the output is written into a system of record, or piped into another automated step, or read later by somebody who has no way to tell it was generated.
| Workflow shape | Value at 85% correct | Cost of a wrong output | Verdict |
|---|---|---|---|
| Draft a reply for a human to send | High — saves the blank page | One edit, caught immediately | Good fit |
| Extract 8 fields from an invoice into a review screen | High — human checks 8 boxes, not 8 documents | Visible in the same screen | Good fit |
| Rank an inbox by likely urgency | High — beats chronological | An item read late | Good fit |
| Auto-send messages to customers | Moderate | Irreversible, invisible until a complaint | Bad fit |
| Write a figure into a report others rely on | Low — nobody can spot the bad one | Silent, compounds downstream | Bad fit |
| Advise a user in a domain they don't know | Low — user cannot verify | Acted on in good faith | Bad fit |
Those two questions rule out a lot of ideas that sound excellent in a roadmap review. Anything that fires an irreversible action — a refund, a cancellation, a deletion, a payment — without a human between the model and the action is on the wrong side of the line, and no amount of prompt work moves it. So is anything whose output lands where the person reading it lacks the context to challenge it.
Long automated chains deserve their own warning. Multi-step plans multiply, they don't average. A pipeline of five steps that are each right 90% of the time finishes correctly about 59% of the time, and a chain of ten lands near 35%. That arithmetic is why so many ambitious agent demos survive a scripted walkthrough and collapse on real traffic. The fix is not a better model; it is fewer steps, checkpoints where a human or a deterministic validator can catch an error before it propagates, and an honest acknowledgement that each unattended hop costs you accuracy.
Design the failure path first
The happy path more or less designs itself. Somebody asks, the model answers, the answer is good, everyone is pleased. The product is decided in the other three states, so design those first — real screens, real copy, before the good case gets a pixel.
Unsure. Abstaining is a feature, and it needs a trigger. The trigger rarely has to be a model confidence score; cheaper signals usually work better. No retrieved document above a similarity threshold. A required field the extractor could not find. An input in a language that isn't in your evaluation set. An input three words long. Route those to a narrower answer, a partial answer with the gaps explicitly marked, or a person. Copy carries a lot of weight here: "I couldn't find anything about refunds in the documents you uploaded" is far better than a bare "I don't know", and both are better than an invention.
Wrong. Assume it happens and make correction the cheapest available action. Present output as a draft to be edited rather than a fact to be accepted. Put the evidence next to the claim so that checking costs a glance instead of a search. Never discard the user's original input — nothing erodes goodwill faster than a bad rewrite that ate the thing it replaced. And capture the correction, because a correction is simultaneously your quality metric and your next evaluation example.
Unavailable. Providers rate-limit, time out, and have bad days. Decide in advance which of three things happens: degrade to something deterministic (keyword search instead of semantic search, a template instead of a generated message), queue the work and notify when it lands, or say plainly that the feature is unavailable and get out of the way. What you must not ship is a spinner with no terminal state, or an empty result that is indistinguishable from "there was nothing to find".
Here is the retention argument for all this, stated bluntly. An abstention costs one interaction. A confident wrong answer costs the user's mental model of your product. Once someone has caught the feature fabricating, they start checking everything it produces — and the checking was exactly the cost your feature existed to remove. The value proposition dies before the feature does.
Trust is a budget that depletes
Every user arrives with a small line of credit. Correct outputs top it up slowly. Visible errors drain it fast, and not linearly: the first bad experience does far more damage than the fifth good one repairs, because early on there is no track record to absorb the hit. A few consequences follow directly.
- Launch narrow. Ship the slice of inputs you are most confident about, even if it feels unambitious. A feature that is excellent on 30% of cases and absent on the rest outperforms one that is mediocre everywhere.
- Watch where first contact happens. If a new user's first encounter with the feature is also their hardest input, you have arranged for the worst possible introduction. Put the entry point where the easy cases live.
- Under-claim in the copy. A panel labelled "suggested draft" survives errors that a panel labelled "answer" does not. The label sets the verification expectation, and the verification expectation decides whether a mistake reads as a normal limitation or as a broken promise.
- Never make the same visible mistake twice to the same person. One error reads as probabilistic. The same error again reads as broken, and that judgement is very hard to reverse.
One tempting move to avoid: printing a confidence percentage in the interface. Users cannot calibrate a number like that, and a wrong answer stamped "82% confident" is more annoying than one with no stamp at all. Express confidence through behaviour instead — hedge the wording, show fewer claims, cite the source, or abstain.
The eval set is a product artefact
An evaluation set is a file of real inputs paired with the behaviour you would accept. It is the most under-owned artefact in AI product work, and it should belong to whoever owns the feature, not to whoever owns the model. The reason is simple: it is the written-down definition of what the feature is for. Every scoping decision from the section above turns into a row in this file, including the decisions about what the feature must refuse.
Left to engineers, the set fills up with inputs that are convenient — clean, short, in one language, of the kind that already works. You need it to be representative, which means it has to include the ugly material:
- The twenty most common real inputs, verbatim, typos and all.
- Every input that has ever embarrassed the feature in a demo or a bug report.
- Out-of-scope and adversarial inputs where the correct behaviour is a refusal or an abstention — with that expected abstention written down as the pass condition.
- Mixed-script and multi-language inputs, plus the extremes of length in both directions.
- Genuinely ambiguous inputs, labelled as ambiguous, where asking a clarifying question is the right answer.
Since you can't write down the one correct output for a generative task, label the checkable properties instead. A row from a real set looks like this — note that the expectations are assertions a script can run, not a paragraph of opinion.
{"id": "eval-041", "input_ref": "ticket/8812", "note": "customer mixes two order numbers, one already refunded", "expect": [ {"kind": "field_equals", "field": "category", "value": "billing"}, {"kind": "field_in", "field": "order_id", "value": ["A-4471", "A-4472"]}, {"kind": "must_mention", "text": "already refunded"}, {"kind": "must_not_promise", "pattern": "we will refund"}, {"kind": "max_sentences", "field": "summary", "value": 3} ], "tags": ["ambiguous_order", "mixed_script"]}Every expectation there is mechanically gradable, which means the set can run on every prompt change without a human in the loop. Rows that need judgement still exist, and those are the ones worth a human review pass; keeping them separate stops the gradable majority from being held hostage to the slow minority.
The set grows from production. Sample real traffic weekly, and turn every complaint into a row. A few dozen rows drawn from actual users beat a thousand synthetic ones, because synthetic inputs inherit the imagination of whoever generated them — which is the same imagination that scoped the feature and therefore misses the same things.
One thing to be honest about: offline scores do not predict product success on their own. A set that improves from 71% to 84% pass rate might change nothing a user notices. Four reasons. The set samples only the inputs you thought of. The grader is not the user. The score ignores latency, cost, and interface entirely. And gains concentrated in the middle of the distribution are invisible to people who only remember the tail. Use the eval set for what it is genuinely good at — catching regressions and comparing two candidates before you expose either to a human. Adoption gets measured in the product.
Metrics that describe what the user did next
Usage counts are seductive because they always go up after a launch and they never tell you anything. Requests served, monthly users of the feature, tokens consumed, "AI interactions" — all of these rise when you add a button to a busy screen. Thumbs-up rates are barely better, since only a small and unrepresentative slice of users ever clicks one.
The metrics that carry information all describe what happened to the output after it appeared.
- Acceptance rate — the fraction of outputs that were used at all: sent, saved, applied, merged, committed. The single most informative number you have.
- Edit fraction — how much of the output changed before use. A continuous quality signal that costs nothing to collect and needs no labelling.
- Correction rate — the fraction where a user fixed a factual error rather than a stylistic one. Worth separating from ordinary editing, because factual corrections are what spend the trust budget.
- Escalation rate — the fraction that ended up with a human anyway. For triage and support features this is the number that decides whether the feature pays for itself.
- Task completion — did the user finish the job they arrived to do, measured the same way you measured it before the feature existed.
- Time-to-outcome — elapsed time end to end, verification included. This is the metric that catches the most common quiet failure: the feature is used, liked, and saves nobody any time.
Getting these requires instrumenting the pair, not the call. Emit one event when an output is produced and another when a user acts on it, joined by a request id, and keep enough of the input to pull failures back into the eval set. Then the numbers are a query rather than a project.
-- ai_outputs: one row per generated output-- ai_output_actions: one row per user action on an outputselect date_trunc('day', o.created_at) as day, count(*) as outputs, avg((coalesce(a.action,'ignored') = 'accepted')::int) as acceptance_rate, avg((coalesce(a.action,'ignored') = 'escalated')::int) as escalation_rate, avg((a.correction_type = 'factual')::int) as correction_rate, avg(a.chars_changed::float / nullif(o.output_chars, 0)) as mean_edit_fraction, percentile_cont(0.5) within group (order by o.latency_ms) as p50_latency_msfrom ai_outputs oleft join ai_output_actions a on a.output_id = o.idwhere o.feature = 'ticket_summary' and o.created_at > now() - interval '30 days'group by 1order by 1;The left join combined with the coalesce is the important detail: an output nobody touched counts as a miss, which is exactly what it is. Reporting acceptance only over outputs that received some action is the most common way teams accidentally flatter themselves by a factor of two or three.
Cost per request against value per accepted output
Here is the arithmetic that catches teams out. You pay per request. You earn per accepted output. If one output in three is accepted, the effective cost of a unit of value is three times the sticker price. Add a retry, a validation pass, and a planning step, and the multiplier compounds:
effective cost per useful output = (calls per attempt × cost per call) ÷ acceptance rate
An agent flow making eight calls per attempt at a 40% acceptance rate is paying roughly twenty times the headline price of a single call. That is survivable if one accepted output is worth ten minutes of a skilled person's time, and fatal if it saves thirty seconds of a cheap task. So name the value side explicitly, in a currency the business already tracks: minutes of handling time, one deflected contact, one fewer manual review. If nobody can state that number even roughly, the feature has no economics — only an invoice.
The levers, in the order they usually pay off: cut calls per attempt; raise the acceptance rate through scoping and interface work, which is nearly always cheaper than reaching for a bigger model; cache aggressively wherever inputs repeat; and trim the context you send. That last one is the quiet cost driver — retrieving twenty documents when three would do inflates the price of every request, and often lowers quality by burying the relevant passage in noise.
Latency belongs in the same conversation, because it does not merely affect satisfaction. It decides which product you are building.
| Response time | Interaction pattern it supports | What it demands |
|---|---|---|
| Under ~300 ms | Inline and speculative — completions, suggestions as you type | Dismissing must be free, so being wrong is nearly costless |
| Roughly 1–3 s | Ask and wait — one request, one visible answer | A clear affordance and a cancel; one request at a time |
| Roughly 5–30 s | No longer conversational — a job with progress | Streaming so the first token arrives fast, or partial results |
| Minutes | Background work the user walks away from | Somewhere for results to live, a notification, idempotent retries |
A twenty-second inline autocomplete is not a slow autocomplete. It is a different feature wearing the wrong interface. Streaming genuinely helps perceived latency and is worth doing, but it does not shorten the time until the user knows whether the output is any good — the verification cost stays exactly where it was.
A spec sheet you can copy
All of the above compresses into one short document per feature. Keep it in the repository next to the code, version it, and make it the thing you argue over in review rather than the thing you reconstruct after launch.
feature: support_ticket_summaryversion: 3intended_use: > Given one support ticket thread, produce a three-sentence summary plus extracted fields so an agent can pick up the ticket without reading it. Read-only. Always shown to an agent before any customer sees anything.out_of_scope: - drafting or sending replies to the customer - threads containing card or health data (routed to the manual queue) - threads longer than 40 messages (truncation changes the meaning) - locales outside en / hi / ta (not represented in the eval set)input_contract: messages: {min: 1, max: 40, required: true} locale: {enum: [en, hi, ta], required: true} max_input_chars: 24000output_schema: summary: {type: string, max_sentences: 3} category: {enum: [billing, delivery, account, other]} order_id: {type: string, nullable: true} evidence: {type: array, items: message_id, min_items: 1} status: {enum: [ok, low_evidence]}fallback: low_evidence: show thread unsummarised, fields marked "not found" schema_invalid: retry once, then show thread unsummarised timeout_ms: 4000, then show thread unsummarised and log input for eval provider_error: banner "summary unavailable", agent workflow unchangedsuccess_metrics: acceptance_rate: ">= 0.60 used without a factual correction" correction_rate: "<= 0.10 sustained over 30 days" time_to_first_reply: "median below the no-summary baseline" cost_per_accepted: "<= one minute of agent handling time"guardrails: never_output: commitments about refunds, credits, or delivery dates human_review: every output — an agent reads it before the customer is contactedeval_set: evals/ticket_summary.v3.jsonlTwo fields do the heaviest lifting. out_of_scope is what you point at when somebody asks whether the feature could also send the reply; it converts a recurring scoping argument into a documented decision with a reason attached. fallback is the field most often left blank, and that blank is why so many AI features have an infinite spinner as their error state. If you fill in nothing else, fill in those two.
Reading the first two weeks
Launch a probabilistic feature and you get a pile of numbers that only mean something in combination. These pairings cover most of what you will actually see.
- Requests high, acceptance low. The output is not shaped like what the job needs, or the feature sits in the wrong place in the flow. Re-scope before touching the model — this is almost never a model problem.
- Acceptance high, time-to-outcome flat. You added a verification step about as expensive as the work you removed. Reduce what has to be checked: shorter outputs, fewer fields, evidence shown inline.
- Acceptance high, time-to-outcome down, escalation up. The feature is winning the easy cases and quietly pushing the hard ones downstream. Check whether difficult cases now arrive later and more annoyed than they used to.
- Eval scores green, usage falling. Your eval set no longer resembles production. Pull fifty real inputs from last week and compare them against the set; the gap is your roadmap.
- Correction rate low overall but concentrated in a handful of users. Those users have the inputs you did not scope for. Read their cases — they are your next eval rows, and possibly your next out_of_scope line.
- Cost per accepted output rising while acceptance is flat. Inputs are getting longer or retries are climbing. Look at calls per attempt and retrieved context size before you look at model pricing.
None of these readings are interpretable unless you know what was live when you took them. Record the model version, the prompt version, and the eval set version alongside every metric snapshot; without that, a movement in acceptance rate is an unattributable mystery and the team ends up arguing from memory. Version those three together, and roll them back together.
There is one signal worth more than everything above, and it arrives without any instrumentation at all. The feature is genuinely working when it goes down for twenty minutes and people complain.