Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
The launch review: a checklist you can reuse
TiffinGo's first launch-review meeting took two and a half hours and ended without a decision. Security asked about prompt injection, finance asked what it would cost at Diwali volume, the support head asked what agents would see during an outage, and legal asked what customer data went to the vendor. Each answer existed, somewhere, in someone's head.
The second meeting took 45 minutes. The difference was a checklist, sent a week earlier, with one owner and one piece of evidence per line. The meeting reviewed the evidence and the few open items, not the whole feature.
A checklist does not make a feature safe. It makes sure that the work of the previous five sections has actually been done and can be shown. Everything on it has appeared earlier in this course.
How to use the checklist
Each item has an owner and needs evidence: a link to a document, a dashboard, a test result or a changelog entry. "Yes, we did that" is not evidence. Items marked blocking must be complete before any real traffic. The others must have an owner and a date.
Send the filled-in checklist to reviewers a week before the meeting. Most questions get answered in comments, and the meeting is left with the genuine disagreements.
Scope and product
| Item | Evidence | Blocking |
|---|---|---|
| One-page scope: in scope, out of scope, level of autonomy | Scope document | Yes |
| Value and risk estimated in rupees, with assumptions | Scope document, value/risk section | No |
| Failure boundaries listed, with the hard ones enforced in code | Boundary tests passing | Yes |
| Review screen shows working, source words and review level | Screenshot and agent pilot feedback | Yes |
| Every human decision logged with release version | Sample log rows | Yes |
Interface and releases
| Item | Evidence | Blocking |
|---|---|---|
| Prompt contract with inputs, outputs, rules and refusals | prompts/order_issue/v6/system.txt | Yes |
| Structured output, three-layer validation, one repair, manual fallback | Validator unit tests | Yes |
| Only required fields sent to the model; no personal data beyond need | render_input review | Yes |
| Releases versioned, hashed, stamped on every draft | Changelog, sample draft | Yes |
| Shadow mode and percentage flag working; rollback tested | Rollback drill record | Yes |
Evaluation
| Item | Evidence | Blocking |
|---|---|---|
| Eval set from real, anonymised tickets, with category table | Eval README | Yes |
| Critical categories pass completely | Latest gate report | Yes |
| Judge calibrated against humans, 90% or better per criterion | Calibration sheet | No |
| Failure taxonomy current for the release | Taxonomy table | Yes |
| Regression gate runs on every prompt and model change | CI configuration | Yes |
| Holdout set run before release | Holdout result | No |
Safety and misuse
| Item | Evidence | Blocking |
|---|---|---|
| Worst case of a fully fooled model written down and bounded | Threat notes | Yes |
| Amounts computed in code; caps and escalations enforced in code | Unit tests | Yes |
| Tools read-only and scoped to the ticket's order | Handler tests | Yes |
| Injection flagging and daily fraud report | Sample report | No |
| Vendor data terms reviewed: retention, training use, region | Legal sign-off | Yes |
Operations and cost
| Item | Evidence | Blocking |
|---|---|---|
| Timeouts, one retry, circuit breaker, fallback release gate-approved | Load test and gate report | Yes |
| Degraded mode tested with the model switched off | Drill record | Yes |
| Dashboards and alerts for the metrics in the drift lesson | Dashboard link | Yes |
| On-call owner and runbook, including weekends | Rota and runbook link | Yes |
| Cost per ticket measured; cost at 2x volume within budget | Cost sheet | Yes |
| Weekly blind sample scheduled with a named grader | Calendar entry | No |
That is 27 items. For TiffinGo, 22 are blocking. A smaller feature will drop some rows; a riskier one, such as one that talks to customers directly, will add rows for output checks and customer-facing wording.
The rollout plan
- Internal pilot — 10 experienced agents, one city, one week. Exit: unchanged approvals above 65%, no critical incidents, agents' feedback addressed.
- One city, all agents — Bengaluru, two weeks. Exit: metrics inside normal ranges for 7 consecutive days; seeded-mistake catch rate above 80%.
- Three cities — one week. Exit: no city more than 5 points below the others on unchanged approvals.
- All cities — with the on-call rota active and the first weekly blind sample done.
Each stage has a written exit condition decided before it starts. Deciding afterwards invites "it is probably fine". And each stage can go back one step with a flag change.
The 30-day review
Launch is a checkpoint, not the finish. Thirty days after full rollout, hold a short review with four questions. Did the value arrive? Compare handling time and refund consistency with the scope's estimates. What did production teach? Read the taxonomy's new groups. What did it cost? Compare real cost with the estimate. What changes next? Pick the next item on the lifecycle loop, often a scope question: which narrow case, if any, has earned a higher level of autonomy?
Check your understanding
0 of 3 answered
1.Why does each checklist item need a link to evidence rather than a yes or no?
2.Why are rollout exit criteria written before each stage starts?
3.The review finds the eval set has 3 Hyderabad tickets while Hyderabad is 18% of volume. Is that blocking?