AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Error analysis: turning failures into a taxonomy


When v3 scored 97 of 120, the first reaction was to open the first failure, see that the model had misread "2 roti kam", and start adding Hinglish instructions to the prompt. That would have fixed perhaps four tickets while the largest problem, worth three times as much money, sat unread further down the list.

Error analysis is the discipline of reading all the failures before fixing any. It is slow, a little boring, and the highest-return hour in the whole evaluation process. It turns a score into a plan.

Nothing in this lesson needs clever tooling. It needs a spreadsheet, the 23 failed drafts, and the patience to read them one by one.

v3's 23 failures, groupedWrong unit price6Model doesmoney mathArithmetic slip3Model doesmoney mathHinglish misread4No examplesCity policy missing4Fact not in promptRestaurant closed3No live dataVeg given meat2Rule too implicitNeedless escalation1Over-cautionFailureCountLikely cause
Nine of 23 failures shared one cause — the model doing money arithmetic — which reading any single failure would never have shown.

Step one: one line per failure

Open each failed row next to its expected output and write one plain sentence describing what went wrong. Do not name a cause yet, and do not group. You are describing, not diagnosing.

Text
t0017  1 of 4 Butter Roti missing; refunded 120, the whole line price, instead of 30t0031  "2 roti kam aaye" read as 2 rotis received; refunded 2 of 4 instead of 2 missingt0044  30% of 185 written as 65 (should be 56)t0058  "paneer ki jagah chicken" drafted as R3 replacement; should escalate food_safetyt0066  Mumbai, July, "bag was wet, rotis soggy"; model used R2 (30%); monsoon rule MUM-M2 gives 50%t0071  drafted "replace" at 11:40 pm; restaurant closed at 11:00t0083  customer wrote "I paid 250 for dal"; model used 250, order says 220...

This takes about two minutes per failure, so under an hour for 23. The one-line habit forces you to look at the actual input and output, not your memory of what the model "usually" does.

Step two: group into a taxonomy

Now read the lines and group similar ones. Name each group by what went wrong, then add a count, an example, a cost and a likely cause. Here is TiffinGo's taxonomy for v3.

Failure typeCountExampleCost when uncaughtLikely cause
Wrong unit price61 of 4 rotis refunded at the ₹120 line price₹30–170 over-refundModel does money arithmetic
Arithmetic slip330% of ₹185 written as ₹65₹5–20Model does money arithmetic
Hinglish misread4"2 roti kam" read as 2 received₹30–60No mixed-language examples
City policy missing4Mumbai monsoon rain-damage rule not applied₹20–60 under-refundFact not in the prompt
Replacement when restaurant closed3"replace" drafted at 11:40 pmAgent redo, customer waitsLive data not in the input
Missed veg-meat escalation2"paneer ki jagah chicken" treated as R3Serious harm to trustRule too implicit, no examples
Needless escalation1Simple missing raita marked "unclear"About 4 minutes of senior timeOver-cautious wording
Total23

Three things stand out that no single failure showed. The two money groups together are 9 of 23, the largest share. Four different groups have causes outside the prompt's wording: missing facts, missing live data, and arithmetic the model should not be doing. And the smallest group by count, veg-meat, is the most serious.

Step three: find the cause, not the symptom

A symptom is what the output looked like. A cause is why. Two failures with the same symptom can have different causes, and the fix follows the cause.

"Wrong amount" was the symptom for 9 tickets. Reading them showed two causes. Six were wrong unit prices, where the model misunderstood what a line's price covered. Three were plain arithmetic slips. Both trace back to one design choice: the model is doing money arithmetic at all. More instructions like "calculate carefully" might reduce the slips, but would not remove them. Moving the arithmetic into code removes the whole group.

Likewise, "city policy missing" is not a prompt-wording problem. The Mumbai monsoon rule is simply not in the contract. No rewording fixes a missing fact.

Step four: prioritise

Rank each group by three things: how often it happens, how much it costs when it does, and how much the fix costs. A simple way is to sort by count × cost, then look at fix effort.

GroupFrequency × costFixFix effort
Money (price + arithmetic)HighModel returns units and rule; code computes amounts1 day
Veg-meat escalationLow count, very high costTwo examples, keyword net extended2 hours
Hinglish misreadMediumThree worked examples in the contract2 hours
City policyMediumRetrieval over the policy (section 5)1 week
Restaurant closedLow–mediumLive status tool (section 5)1 week
Needless escalationLowWatch; do not fix yetNone

The cheap, high-value fixes go first, into v4. The expensive ones are scheduled, and their need is now documented with evidence. One group is deliberately left alone: a single needless escalation costs four minutes, and "fixing" it risks making the model less cautious everywhere.

Keep the taxonomy alive

The taxonomy is not a one-time report. Recount it after every release, and extend it with failures from production: every rejected draft and every heavy edit an agent makes is a candidate. New groups will appear. When TiffinGo launched in Pune, a new group, "packaging charge refunded as an item", appeared within a week.

A taxonomy that never changes usually means nobody is reading failures any more.

Check your understanding

0 of 3 answered

1.The first failure you open is a Hinglish misread. What should you do next?

2.Nine failures show a wrong refund amount. Six used the wrong unit price and three were arithmetic slips. What is the root-cause fix?

3.The veg-meat group has only 2 failures. Why is it fixed first anyway?