Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Reading your failures: which capability fixes which failure
With the v3 taxonomy on the wall, the TiffinGo team held a planning meeting. Within ten minutes there were four proposals: "add RAG", "use the biggest model", "fine-tune on our tickets", and "make it an agent". Each had a champion and a blog post behind it. None of them was connected to a specific row in the taxonomy.
Capabilities are not upgrades. Each one fixes a particular kind of failure and adds its own cost, latency and new ways to fail. Retrieval adds a search step that can miss. Tools add calls to other services that can time out. A larger model adds cost on every ticket, including the 80% that were already right.
This lesson gives you a map from the cause of a failure to the capability that fixes it, and an order to try them in. Then it applies the cheapest fix to TiffinGo's biggest failure group.
Match the cause to the capability
Every row in a taxonomy has a likely cause. The cause, not the symptom, picks the fix.
| Cause of the failure | Capability that fixes it | TiffinGo example |
|---|---|---|
| The right data exists but is not in the input | Add it to the input in code | paid_inr missing in the first prompt |
| The model is doing work code does exactly | Move that work into code | Money arithmetic, date differences |
| The model misreads a kind of input | Clearer contract, worked examples | Hinglish quantities, veg-meat cases |
| The model lacks facts that are too many or change too often for the prompt | Retrieval | City policy rules, monthly promotions |
| The model needs live data, or must trigger an action | Tool calling | Is the restaurant open right now? |
| The task needs more reasoning than the model has | A more capable model or more reasoning effort | Rare in TiffinGo's taxonomy |
| The model's behaviour is right but too costly or slow at scale | Fine-tuning a smaller model | Not yet |
Notice how few rows point at the model itself. In most real taxonomies, the majority of failures come from missing input, work misplaced in the model, or missing facts. Model capability is usually the smallest group, which is why "use the biggest model" rarely fixes as much as people hope.
Try the cheapest fix first
- Fix the input — hours; nothing new to operate. Add the missing field or remove confusing ones.
- Move work into code — a day; code is exact and testable. Arithmetic, lookups, date logic.
- Improve the contract — hours; costs a few tokens per call. Clearer rules, a few worked examples.
- Add retrieval or a tool — a week or more; adds a service to run and a new failure mode.
- Change the model — hours to try, but changes cost and latency on every ticket.
- Fine-tune — weeks; adds training data, a training pipeline and a model to host and version.
The order is not a law. If the taxonomy says 60% of failures need live data, you go to step 4 quickly. But you should be able to explain why each cheaper step does not fix the group in front of you.
Applying it: money moves into code
The largest group in the v3 taxonomy was money: 6 wrong unit prices and 3 arithmetic slips. The cause for both was that the model was computing amounts. Step 2 applies: the model should decide what happened, and code should decide what it costs.
In contract v4, each line has only item, units, issue and rule. The model no longer returns amount_inr or refund_inr. Code computes them from the order and a pricing table.
1import math23# Pricing per rule, as data. Retrieval later adds city rules to this table.4RULES = {5 "R1": {"percent": 100, "min_inr": 0}, # missing units6 "R2": {"percent": 30, "min_inr": 30}, # cold or badly packed units7 "R3": {"percent": 100, "min_inr": 0}, # wrong item, refunded8}910def rupees(amount: float) -> int:11 return math.floor(amount + 0.5) # halves round up, as the policy says1213def line_amount(line: dict, item: dict, rules: dict) -> int:14 rule = rules[line["rule"]]15 unit_price = item["paid_inr"] / item["qty"]16 return max(rule["min_inr"], rupees(unit_price * line["units"] * rule["percent"] / 100))1718def price_draft(draft: dict, order: dict, rules: dict = RULES) -> dict:19 items = {item["name"]: item for item in order["items"]}20 lines = [{**line, "amount_inr": line_amount(line, items[line["item"]], rules)}21 for line in draft["lines"]]22 total = sum(line["amount_inr"] for line in lines) if draft["action"] == "refund" else 023 return {**draft, "lines": lines, "refund_inr": total}On the cold-dal ticket, the model returns one Butter Roti under R1 and one Dal Makhani under R2. Code computes ₹120 / 4 × 1 = ₹30 and max(₹30, 30% of ₹220) = ₹66, a total of ₹96, the same every time.
rupees exists because Python's built-in round rounds halves to the nearest even number: round(54.5) is 54. The policy says halves round up, so 30% of ₹185, which is ₹55.50, must be ₹56. That is the kind of detail a model will get wrong occasionally and code will get right always, once someone has written it down.
The pricing table is data, not code branches. That choice pays off in the next lesson, when city-specific rules arrive from retrieval with their own percentages.
Examples for the misreads
Two smaller groups also sat at cheap steps. Hinglish misreads (4) and missed veg-meat escalations (2) were cases of the model misreading a kind of input. Step 3 applies: the team added five short worked examples to the contract, three Hinglish tickets with quantities and two veg-meat tickets, each with the correct draft. They also added "chicken", "mutton", "egg", "fish" and "non-veg" to the keyword net, active only when every item in the order is marked vegetarian in the menu data.
Examples are powerful and a little dangerous. The model copies their patterns, including patterns you did not intend. One early example used "2 roti kam" with an order of 4, and the model began assuming 4 rotis in unrelated orders. Keep examples varied, and let the eval set tell you whether they helped.
Check your understanding
0 of 3 answered
1.Six failures come from the model refunding a whole line's price for one missing unit. Which fix addresses the cause?
2.Why is the pricing table written as data (percent and minimum per rule) instead of if-statements?
3.Most of a taxonomy's failures come from missing input data and misplaced arithmetic. What does that suggest about "use the biggest model"?