AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Watching quality drift, cost and latency after launch


Six weeks after launch, with no new release, the share of drafts approved unchanged slipped from 74% to 68% over nine days. No errors, no alerts from the provider, no complaints from the support floor. Agents had simply started editing more drafts.

The cause turned out to be outside the assistant entirely. TiffinGo had launched "combo meals": one order line, such as "Thali Combo", containing dal, sabzi, rice, 3 rotis and raita. Complaints like "combo mein raita nahi tha" (the raita in the combo was missing) had no way to be expressed in the contract, which only knew whole lines. The model did its best, and agents fixed the rest.

This is drift: quality changing without any change to your code, because the world your inputs come from has changed. New menu formats, new cities, festival seasons, policy changes and silent provider updates all cause it. You cannot prevent drift. You can notice it in days instead of months.

Approved-unchanged rate by segment, week of the dip767573783101234missingcoldwrong itemlatecomboordersOverall fell only from 74% to 68%.
Drift hid in one segment that fell forty points; only a query broken down by category showed where to read tickets.

What to watch

Four groups of numbers, each with a normal range and an alert.

MetricWhat it tells youNormal at TiffinGoAlert when
Approved unchangedOverall usefulness of drafts72–76%Below 70% for 2 days
Average refund gap (final minus draft)Drift towards over- or under-refunding₹0 to +₹3Beyond ±₹8 for 2 days
Rejected draftsDrafts agents threw away8–10%Above 13%
Validator rejectionsContract confusion or input changes1–2%Above 3%
Escalation rate, food-safety countBoundaries firing as designed10–14%, about 25 a dayOutside range, either way
Cost per ticket, tokens per ticketPrompt growth, retries, runaway tools₹0.60–0.65Above ₹0.75
p95 time to draft readyProvider or tool slowness5–7 sAbove 15 s
Fallback and degraded shareProvider healthUnder 1%Above 5% in an hour
Injection flagsAttack wavesAbout 0.6%Above 2% in a day

Note the escalation alert fires in both directions. A sudden drop in food-safety escalations is as worrying as a rise: it may mean the model has stopped recognising them.

Human decisions are your production eval, with care

The approval log from section 2 is the richest quality signal you have: every ticket, labelled for free. But it is biased. Agents approve some wrong drafts (automation bias), and they edit some right ones out of habit. A falling approval rate is a reliable alarm. The approval rate itself is not an accuracy number.

So TiffinGo adds one small, unbiased measurement: every week, a senior agent grades 50 randomly sampled drafts against the eval set's rubric, without seeing what the original agent did. It takes about an hour. It gives a production accuracy estimate within about 8 points, catches problems approvals miss, and supplies fresh eval rows.

Drift hides in segments

A 6-point fall in the overall approval rate can be a 40-point fall in one segment. Always break metrics down by complaint category, city, language, release version and, for TiffinGo, restaurant chain.

SQL
SELECT date_trunc('week', created_at)                                    AS week,       category,       count(*)                                                          AS drafts,       round(100.0 * avg((agent_outcome = 'approved_unchanged')::int), 1) AS unchanged_pct,       round(avg(final_refund_inr - draft_refund_inr), 1)                AS avg_refund_gap_inrFROM draft_decisionsWHERE created_at >= now() - interval '8 weeks'GROUP BY 1, 2ORDER BY 1, 2;

Run by category, this query showed the combo problem at once: "missing item" tickets on combo orders had fallen to 31% unchanged, while every other segment was steady. The average gap on those tickets was minus ₹22: drafts were under-refunding, and agents were correcting upwards.

Watch cost like a quality metric

Cost drifts too, and usually for boring reasons. Two weeks after the combo launch, cost per ticket jumped from ₹0.62 to ₹0.91 overnight. The menu service had started returning every combo's sub-items, with descriptions, inside the order record, and render_input passed them all through. Input tokens per ticket went from about 3,600 to 5,250.

The alert fired the next morning. The fix was one line: pass sub-item names and quantities, not descriptions. Without the alert, the extra ₹350 a day would have run for months, because nothing else was wrong. Track tokens per ticket as well as rupees, so a price change and a prompt change are not confused.

From alert to action: a short runbook

  1. What changed on our side? Check releases, config and flags in the window. Roll back first if a release lines up with the change.
  2. What changed in the inputs? Menu formats, new cities or restaurants, policy updates, a festival, a marketing campaign.
  3. What changed at the provider? Model version notes, latency and error dashboards.
  4. Which segment moved? Run the segmented query; read 20 affected tickets end to end.
  5. Capture and fix — add representative tickets to the eval set, update the taxonomy, and fix through the normal release path with the gate.

The last step is the one teams skip. An incident fixed without new eval rows can come back in the next release without anyone noticing.

Check your understanding

0 of 3 answered

1.The unchanged-approval rate is steady overall at 73%, but food-safety escalations fell from about 25 a day to 9. What should happen?

2.Why does TiffinGo add a weekly blind sample of 50 drafts when it already has every agent's decision?

3.Cost per ticket rises from ₹0.62 to ₹0.91 overnight with no release. What is the most likely place to look first?