Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Error analysis: turning failures into a taxonomy
When v3 scored 97 of 120, the first reaction was to open the first failure, see that the model had misread "2 roti kam", and start adding Hinglish instructions to the prompt. That would have fixed perhaps four tickets while the largest problem, worth three times as much money, sat unread further down the list.
Error analysis is the discipline of reading all the failures before fixing any. It is slow, a little boring, and the highest-return hour in the whole evaluation process. It turns a score into a plan.
Nothing in this lesson needs clever tooling. It needs a spreadsheet, the 23 failed drafts, and the patience to read them one by one.
Step one: one line per failure
Open each failed row next to its expected output and write one plain sentence describing what went wrong. Do not name a cause yet, and do not group. You are describing, not diagnosing.
t0017 1 of 4 Butter Roti missing; refunded 120, the whole line price, instead of 30t0031 "2 roti kam aaye" read as 2 rotis received; refunded 2 of 4 instead of 2 missingt0044 30% of 185 written as 65 (should be 56)t0058 "paneer ki jagah chicken" drafted as R3 replacement; should escalate food_safetyt0066 Mumbai, July, "bag was wet, rotis soggy"; model used R2 (30%); monsoon rule MUM-M2 gives 50%t0071 drafted "replace" at 11:40 pm; restaurant closed at 11:00t0083 customer wrote "I paid 250 for dal"; model used 250, order says 220...This takes about two minutes per failure, so under an hour for 23. The one-line habit forces you to look at the actual input and output, not your memory of what the model "usually" does.
Step two: group into a taxonomy
Now read the lines and group similar ones. Name each group by what went wrong, then add a count, an example, a cost and a likely cause. Here is TiffinGo's taxonomy for v3.
| Failure type | Count | Example | Cost when uncaught | Likely cause |
|---|---|---|---|---|
| Wrong unit price | 6 | 1 of 4 rotis refunded at the ₹120 line price | ₹30–170 over-refund | Model does money arithmetic |
| Arithmetic slip | 3 | 30% of ₹185 written as ₹65 | ₹5–20 | Model does money arithmetic |
| Hinglish misread | 4 | "2 roti kam" read as 2 received | ₹30–60 | No mixed-language examples |
| City policy missing | 4 | Mumbai monsoon rain-damage rule not applied | ₹20–60 under-refund | Fact not in the prompt |
| Replacement when restaurant closed | 3 | "replace" drafted at 11:40 pm | Agent redo, customer waits | Live data not in the input |
| Missed veg-meat escalation | 2 | "paneer ki jagah chicken" treated as R3 | Serious harm to trust | Rule too implicit, no examples |
| Needless escalation | 1 | Simple missing raita marked "unclear" | About 4 minutes of senior time | Over-cautious wording |
| Total | 23 |
Three things stand out that no single failure showed. The two money groups together are 9 of 23, the largest share. Four different groups have causes outside the prompt's wording: missing facts, missing live data, and arithmetic the model should not be doing. And the smallest group by count, veg-meat, is the most serious.
Step three: find the cause, not the symptom
A symptom is what the output looked like. A cause is why. Two failures with the same symptom can have different causes, and the fix follows the cause.
"Wrong amount" was the symptom for 9 tickets. Reading them showed two causes. Six were wrong unit prices, where the model misunderstood what a line's price covered. Three were plain arithmetic slips. Both trace back to one design choice: the model is doing money arithmetic at all. More instructions like "calculate carefully" might reduce the slips, but would not remove them. Moving the arithmetic into code removes the whole group.
Likewise, "city policy missing" is not a prompt-wording problem. The Mumbai monsoon rule is simply not in the contract. No rewording fixes a missing fact.
Step four: prioritise
Rank each group by three things: how often it happens, how much it costs when it does, and how much the fix costs. A simple way is to sort by count × cost, then look at fix effort.
| Group | Frequency × cost | Fix | Fix effort |
|---|---|---|---|
| Money (price + arithmetic) | High | Model returns units and rule; code computes amounts | 1 day |
| Veg-meat escalation | Low count, very high cost | Two examples, keyword net extended | 2 hours |
| Hinglish misread | Medium | Three worked examples in the contract | 2 hours |
| City policy | Medium | Retrieval over the policy (section 5) | 1 week |
| Restaurant closed | Low–medium | Live status tool (section 5) | 1 week |
| Needless escalation | Low | Watch; do not fix yet | None |
The cheap, high-value fixes go first, into v4. The expensive ones are scheduled, and their need is now documented with evidence. One group is deliberately left alone: a single needless escalation costs four minutes, and "fixing" it risks making the model less cautious everywhere.
Keep the taxonomy alive
The taxonomy is not a one-time report. Recount it after every release, and extend it with failures from production: every rejected draft and every heavy edit an agent makes is a candidate. New groups will appear. When TiffinGo launched in Pune, a new group, "packaging charge refunded as an item", appeared within a week.
A taxonomy that never changes usually means nobody is reading failures any more.
Check your understanding
0 of 3 answered
1.The first failure you open is a Hinglish misread. What should you do next?
2.Nine failures show a wrong refund amount. Six used the wrong unit price and three were arithmetic slips. What is the root-cause fix?
3.The veg-meat group has only 2 failures. Why is it fixed first anyway?