AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

The launch review: a checklist you can reuse


TiffinGo's first launch-review meeting took two and a half hours and ended without a decision. Security asked about prompt injection, finance asked what it would cost at Diwali volume, the support head asked what agents would see during an outage, and legal asked what customer data went to the vendor. Each answer existed, somewhere, in someone's head.

The second meeting took 45 minutes. The difference was a checklist, sent a week earlier, with one owner and one piece of evidence per line. The meeting reviewed the evidence and the few open items, not the whole feature.

A checklist does not make a feature safe. It makes sure that the work of the previous five sections has actually been done and can be shown. Everything on it has appeared earlier in this course.

27 checklist items, each with a link to evidenceLaunch reviewScope and product — 5Interface and releases — 5Evaluation — 6Safety and misuse — 5Operations and cost — 6
Twenty-two items block launch; the review found two gaps — no weekend on-call and too few Hyderabad tickets — before customers did.

How to use the checklist

Each item has an owner and needs evidence: a link to a document, a dashboard, a test result or a changelog entry. "Yes, we did that" is not evidence. Items marked blocking must be complete before any real traffic. The others must have an owner and a date.

Send the filled-in checklist to reviewers a week before the meeting. Most questions get answered in comments, and the meeting is left with the genuine disagreements.

Scope and product

ItemEvidenceBlocking
One-page scope: in scope, out of scope, level of autonomyScope documentYes
Value and risk estimated in rupees, with assumptionsScope document, value/risk sectionNo
Failure boundaries listed, with the hard ones enforced in codeBoundary tests passingYes
Review screen shows working, source words and review levelScreenshot and agent pilot feedbackYes
Every human decision logged with release versionSample log rowsYes

Interface and releases

ItemEvidenceBlocking
Prompt contract with inputs, outputs, rules and refusalsprompts/order_issue/v6/system.txtYes
Structured output, three-layer validation, one repair, manual fallbackValidator unit testsYes
Only required fields sent to the model; no personal data beyond needrender_input reviewYes
Releases versioned, hashed, stamped on every draftChangelog, sample draftYes
Shadow mode and percentage flag working; rollback testedRollback drill recordYes

Evaluation

ItemEvidenceBlocking
Eval set from real, anonymised tickets, with category tableEval READMEYes
Critical categories pass completelyLatest gate reportYes
Judge calibrated against humans, 90% or better per criterionCalibration sheetNo
Failure taxonomy current for the releaseTaxonomy tableYes
Regression gate runs on every prompt and model changeCI configurationYes
Holdout set run before releaseHoldout resultNo

Safety and misuse

ItemEvidenceBlocking
Worst case of a fully fooled model written down and boundedThreat notesYes
Amounts computed in code; caps and escalations enforced in codeUnit testsYes
Tools read-only and scoped to the ticket's orderHandler testsYes
Injection flagging and daily fraud reportSample reportNo
Vendor data terms reviewed: retention, training use, regionLegal sign-offYes

Operations and cost

ItemEvidenceBlocking
Timeouts, one retry, circuit breaker, fallback release gate-approvedLoad test and gate reportYes
Degraded mode tested with the model switched offDrill recordYes
Dashboards and alerts for the metrics in the drift lessonDashboard linkYes
On-call owner and runbook, including weekendsRota and runbook linkYes
Cost per ticket measured; cost at 2x volume within budgetCost sheetYes
Weekly blind sample scheduled with a named graderCalendar entryNo

That is 27 items. For TiffinGo, 22 are blocking. A smaller feature will drop some rows; a riskier one, such as one that talks to customers directly, will add rows for output checks and customer-facing wording.

The rollout plan

  1. Internal pilot — 10 experienced agents, one city, one week. Exit: unchanged approvals above 65%, no critical incidents, agents' feedback addressed.
  2. One city, all agents — Bengaluru, two weeks. Exit: metrics inside normal ranges for 7 consecutive days; seeded-mistake catch rate above 80%.
  3. Three cities — one week. Exit: no city more than 5 points below the others on unchanged approvals.
  4. All cities — with the on-call rota active and the first weekly blind sample done.

Each stage has a written exit condition decided before it starts. Deciding afterwards invites "it is probably fine". And each stage can go back one step with a flag change.

The 30-day review

Launch is a checkpoint, not the finish. Thirty days after full rollout, hold a short review with four questions. Did the value arrive? Compare handling time and refund consistency with the scope's estimates. What did production teach? Read the taxonomy's new groups. What did it cost? Compare real cost with the estimate. What changes next? Pick the next item on the lifecycle loop, often a scope question: which narrow case, if any, has earned a higher level of autonomy?

Check your understanding

0 of 3 answered

1.Why does each checklist item need a link to evidence rather than a yes or no?

2.Why are rollout exit criteria written before each stage starts?

3.The review finds the eval set has 3 Hyderabad tickets while Hyderabad is 18% of volume. Is that blocking?