Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Designing the product around uncertain output (drafts, confidence, human in the loop)
In the third week of the pilot, TiffinGo's drafts were approved 96% of the time, with a median review of nine seconds. The team celebrated, then ran a test. They added 20 drafts with deliberate mistakes into the queue, such as a wrong item or a doubled amount. Agents caught 4 of them.
The human in the loop had stopped being in the loop. This is automation bias: when a system is usually right, people stop checking it. It is not laziness. It is what any person does when a tool earns trust and the queue is long.
So "a human reviews every draft" is not a safety design by itself. The product has to make review fast when the draft is easy, slow when the draft is risky, and honest about what the system does not know.
Choose the level of autonomy
AI features sit on a ladder. Each step up saves more human time and removes a safety check.
| Level | What the system does | TiffinGo example | When it fits |
|---|---|---|---|
| Suggest | Shows options; the human decides and acts | "Possible rules: R1, R2" | Low accuracy, or a new feature |
| Draft | Prepares a full decision; the human approves or edits | The Order Issue Assistant | Good accuracy, costly errors |
| Act with undo | Acts, and a human can reverse it for a time | Auto-refund under ₹50, reversible for 24 hours | Very high accuracy, small reversible errors |
| Act silently | Acts, with only sampled review | Complaint tags for analytics | Errors are cheap and average out |
TiffinGo chose Draft. Suggest wastes the assistant's accuracy: agents still do most of the work. Act with undo is tempting for small refunds, but money sent to a customer is awkward to take back, and the team had no production data yet. The ladder is not permanent. A narrow slice can move up once months of data prove it.
Make review fast and real
A good review screen does two things at once. It lets the agent confirm an easy draft in seconds, and it forces attention where attention is needed.
Show the working, not only the answer. One line per item, the rule behind it, and the numbers used. "Refund ₹96" asks for trust. "1 of 4 Butter Roti, ₹30 (R1); Dal Makhani cold, 30% of ₹220 = ₹66 (R2)" can be checked in five seconds against the order on the same screen.
Put the source next to the claim. Show the customer's exact words beside the draft, with the words the model relied on highlighted. If the draft says "roti missing" and nothing in the ticket mentions roti, the agent sees it at once.
Add friction only where risk is high. A green draft gets a one-click Approve. An amber draft needs the agent to tick each line. A red draft is not shown at all; the agent works the ticket by hand. Friction everywhere would be ignored everywhere.
Keep measuring attention. The 20 seeded mistakes became a standing practice: two known-wrong drafts per agent per week, reviewed with the agent afterwards as coaching, never as punishment. The catch rate rose from 20% to 85% in a month.
Do not trust the model's confidence number
A tempting design is to add "confidence": 0.93 to the output and show it. Resist it. A model's stated confidence is text it generates, not a measured probability. In TiffinGo's pilot, drafts with a self-reported confidence above 0.9 were wrong 7% of the time, and drafts below 0.9 were wrong 11%. That difference is too small to act on.
Use signals you can compute and test instead:
1from dataclasses import dataclass23@dataclass4class Checks:5 schema_ok: bool # the draft parsed and matched the schema6 items_in_order: bool # every affected item exists in the order7 refund_within_paid: bool # refund is not more than the affected items cost8 safety_keyword_hit: bool # the keyword net fired on the ticket text910def review_level(draft: dict, checks: Checks) -> str:11 if not (checks.schema_ok and checks.items_in_order):12 return "red" # hide the draft; the agent works the ticket by hand13 if checks.safety_keyword_hit and draft["action"] != "escalate":14 return "red"15 if not checks.refund_within_paid:16 return "red"17 if draft["action"] == "escalate" or draft["refund_inr"] >= 300:18 return "amber" # tick each line before approving19 return "green" # one-click approveEach input is a fact the code checked, not an opinion the model offered. You can test this function with ordinary unit tests, and you can check on the eval set how often each level is wrong. In TiffinGo's data, green drafts are wrong about 3% of the time, amber about 12%, and red is shown for about 4% of tickets.
A stronger signal is agreement: run the model twice and compare. If both runs give the same action and amount, the draft is more likely right. This doubles the cost of the call, so use it only for amber tickets, where it pays for itself.
Self-reported confidence
- Generated text, not a measurement
- Barely separates right from wrong drafts
- Cannot be unit-tested
- Tempts agents to skip review when high
Computed signals
- Facts checked by code: schema, items, amounts, keywords
- Measurably separates green from amber from red
- Tested like any other function
- Drives friction where risk is real
Capture every human decision
Every approval, edit and rejection is a label you did not have to pay for. Log it with everything needed to replay the decision.
1{2 "ticket_id": 88121,3 "prompt_version": "v3",4 "model": "configured-model-id",5 "review_level": "green",6 "draft": {"action": "refund", "refund_inr": 96, "rules": ["R1", "R2"]},7 "final": {"action": "refund", "refund_inr": 96},8 "agent_outcome": "approved_unchanged",9 "review_seconds": 4110}This record feeds three later lessons. Edited and rejected drafts become new eval rows. The share of unchanged approvals becomes the main quality signal in production. The gap between draft and final amounts shows whether the assistant drifts towards over- or under-refunding.
Check your understanding
0 of 3 answered
1.Agents approve 96% of drafts in a median of nine seconds, and catch only 4 of 20 seeded errors. What is the main problem?
2.Why does TiffinGo avoid showing the model's self-reported confidence?
3.Which is the best reason to log every agent decision with the prompt version?