Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Watching quality drift, cost and latency after launch
Six weeks after launch, with no new release, the share of drafts approved unchanged slipped from 74% to 68% over nine days. No errors, no alerts from the provider, no complaints from the support floor. Agents had simply started editing more drafts.
The cause turned out to be outside the assistant entirely. TiffinGo had launched "combo meals": one order line, such as "Thali Combo", containing dal, sabzi, rice, 3 rotis and raita. Complaints like "combo mein raita nahi tha" (the raita in the combo was missing) had no way to be expressed in the contract, which only knew whole lines. The model did its best, and agents fixed the rest.
This is drift: quality changing without any change to your code, because the world your inputs come from has changed. New menu formats, new cities, festival seasons, policy changes and silent provider updates all cause it. You cannot prevent drift. You can notice it in days instead of months.
What to watch
Four groups of numbers, each with a normal range and an alert.
| Metric | What it tells you | Normal at TiffinGo | Alert when |
|---|---|---|---|
| Approved unchanged | Overall usefulness of drafts | 72–76% | Below 70% for 2 days |
| Average refund gap (final minus draft) | Drift towards over- or under-refunding | ₹0 to +₹3 | Beyond ±₹8 for 2 days |
| Rejected drafts | Drafts agents threw away | 8–10% | Above 13% |
| Validator rejections | Contract confusion or input changes | 1–2% | Above 3% |
| Escalation rate, food-safety count | Boundaries firing as designed | 10–14%, about 25 a day | Outside range, either way |
| Cost per ticket, tokens per ticket | Prompt growth, retries, runaway tools | ₹0.60–0.65 | Above ₹0.75 |
| p95 time to draft ready | Provider or tool slowness | 5–7 s | Above 15 s |
| Fallback and degraded share | Provider health | Under 1% | Above 5% in an hour |
| Injection flags | Attack waves | About 0.6% | Above 2% in a day |
Note the escalation alert fires in both directions. A sudden drop in food-safety escalations is as worrying as a rise: it may mean the model has stopped recognising them.
Human decisions are your production eval, with care
The approval log from section 2 is the richest quality signal you have: every ticket, labelled for free. But it is biased. Agents approve some wrong drafts (automation bias), and they edit some right ones out of habit. A falling approval rate is a reliable alarm. The approval rate itself is not an accuracy number.
So TiffinGo adds one small, unbiased measurement: every week, a senior agent grades 50 randomly sampled drafts against the eval set's rubric, without seeing what the original agent did. It takes about an hour. It gives a production accuracy estimate within about 8 points, catches problems approvals miss, and supplies fresh eval rows.
Drift hides in segments
A 6-point fall in the overall approval rate can be a 40-point fall in one segment. Always break metrics down by complaint category, city, language, release version and, for TiffinGo, restaurant chain.
1SELECT date_trunc('week', created_at) AS week,2 category,3 count(*) AS drafts,4 round(100.0 * avg((agent_outcome = 'approved_unchanged')::int), 1) AS unchanged_pct,5 round(avg(final_refund_inr - draft_refund_inr), 1) AS avg_refund_gap_inr6FROM draft_decisions7WHERE created_at >= now() - interval '8 weeks'8GROUP BY 1, 29ORDER BY 1, 2;Run by category, this query showed the combo problem at once: "missing item" tickets on combo orders had fallen to 31% unchanged, while every other segment was steady. The average gap on those tickets was minus ₹22: drafts were under-refunding, and agents were correcting upwards.
Watch cost like a quality metric
Cost drifts too, and usually for boring reasons. Two weeks after the combo launch, cost per ticket jumped from ₹0.62 to ₹0.91 overnight. The menu service had started returning every combo's sub-items, with descriptions, inside the order record, and render_input passed them all through. Input tokens per ticket went from about 3,600 to 5,250.
The alert fired the next morning. The fix was one line: pass sub-item names and quantities, not descriptions. Without the alert, the extra ₹350 a day would have run for months, because nothing else was wrong. Track tokens per ticket as well as rupees, so a price change and a prompt change are not confused.
From alert to action: a short runbook
- What changed on our side? Check releases, config and flags in the window. Roll back first if a release lines up with the change.
- What changed in the inputs? Menu formats, new cities or restaurants, policy updates, a festival, a marketing campaign.
- What changed at the provider? Model version notes, latency and error dashboards.
- Which segment moved? Run the segmented query; read 20 affected tickets end to end.
- Capture and fix — add representative tickets to the eval set, update the taxonomy, and fix through the normal release path with the gate.
The last step is the one teams skip. An incident fixed without new eval rows can come back in the next release without anyone noticing.
Check your understanding
0 of 3 answered
1.The unchanged-approval rate is steady overall at 73%, but food-safety escalations fell from about 25 a day to 9. What should happen?
2.Why does TiffinGo add a weekly blind sample of 50 drafts when it already has every agent's decision?
3.Cost per ticket rises from ₹0.62 to ₹0.91 overnight with no release. What is the most likely place to look first?