AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Reading your failures: which capability fixes which failure


With the v3 taxonomy on the wall, the TiffinGo team held a planning meeting. Within ten minutes there were four proposals: "add RAG", "use the biggest model", "fine-tune on our tickets", and "make it an agent". Each had a champion and a blog post behind it. None of them was connected to a specific row in the taxonomy.

Capabilities are not upgrades. Each one fixes a particular kind of failure and adds its own cost, latency and new ways to fail. Retrieval adds a search step that can miss. Tools add calls to other services that can time out. A larger model adds cost on every ticket, including the 80% that were already right.

This lesson gives you a map from the cause of a failure to the capability that fixes it, and an order to try them in. Then it applies the cheapest fix to TiffinGo's biggest failure group.

The fix ladder, cheapest rung at the bottomFix the inputMove workinto codeImprovethe contractAddretrieval or a toolChange the modelFine-tunetopbottomMoney errors went from 9 to 0 on the second rung.
Each rung up adds cost and a new way to fail, so the cause of a failure — not fashion — decides how high you climb.

Match the cause to the capability

Every row in a taxonomy has a likely cause. The cause, not the symptom, picks the fix.

Cause of the failureCapability that fixes itTiffinGo example
The right data exists but is not in the inputAdd it to the input in codepaid_inr missing in the first prompt
The model is doing work code does exactlyMove that work into codeMoney arithmetic, date differences
The model misreads a kind of inputClearer contract, worked examplesHinglish quantities, veg-meat cases
The model lacks facts that are too many or change too often for the promptRetrievalCity policy rules, monthly promotions
The model needs live data, or must trigger an actionTool callingIs the restaurant open right now?
The task needs more reasoning than the model hasA more capable model or more reasoning effortRare in TiffinGo's taxonomy
The model's behaviour is right but too costly or slow at scaleFine-tuning a smaller modelNot yet

Notice how few rows point at the model itself. In most real taxonomies, the majority of failures come from missing input, work misplaced in the model, or missing facts. Model capability is usually the smallest group, which is why "use the biggest model" rarely fixes as much as people hope.

Try the cheapest fix first

  1. Fix the input — hours; nothing new to operate. Add the missing field or remove confusing ones.
  2. Move work into code — a day; code is exact and testable. Arithmetic, lookups, date logic.
  3. Improve the contract — hours; costs a few tokens per call. Clearer rules, a few worked examples.
  4. Add retrieval or a tool — a week or more; adds a service to run and a new failure mode.
  5. Change the model — hours to try, but changes cost and latency on every ticket.
  6. Fine-tune — weeks; adds training data, a training pipeline and a model to host and version.

The order is not a law. If the taxonomy says 60% of failures need live data, you go to step 4 quickly. But you should be able to explain why each cheaper step does not fix the group in front of you.

Applying it: money moves into code

The largest group in the v3 taxonomy was money: 6 wrong unit prices and 3 arithmetic slips. The cause for both was that the model was computing amounts. Step 2 applies: the model should decide what happened, and code should decide what it costs.

In contract v4, each line has only item, units, issue and rule. The model no longer returns amount_inr or refund_inr. Code computes them from the order and a pricing table.

Python
import math# Pricing per rule, as data. Retrieval later adds city rules to this table.RULES = {    "R1": {"percent": 100, "min_inr": 0},   # missing units    "R2": {"percent": 30, "min_inr": 30},   # cold or badly packed units    "R3": {"percent": 100, "min_inr": 0},   # wrong item, refunded}def rupees(amount: float) -> int:    return math.floor(amount + 0.5)  # halves round up, as the policy saysdef line_amount(line: dict, item: dict, rules: dict) -> int:    rule = rules[line["rule"]]    unit_price = item["paid_inr"] / item["qty"]    return max(rule["min_inr"], rupees(unit_price * line["units"] * rule["percent"] / 100))def price_draft(draft: dict, order: dict, rules: dict = RULES) -> dict:    items = {item["name"]: item for item in order["items"]}    lines = [{**line, "amount_inr": line_amount(line, items[line["item"]], rules)}             for line in draft["lines"]]    total = sum(line["amount_inr"] for line in lines) if draft["action"] == "refund" else 0    return {**draft, "lines": lines, "refund_inr": total}

On the cold-dal ticket, the model returns one Butter Roti under R1 and one Dal Makhani under R2. Code computes ₹120 / 4 × 1 = ₹30 and max(₹30, 30% of ₹220) = ₹66, a total of ₹96, the same every time.

rupees exists because Python's built-in round rounds halves to the nearest even number: round(54.5) is 54. The policy says halves round up, so 30% of ₹185, which is ₹55.50, must be ₹56. That is the kind of detail a model will get wrong occasionally and code will get right always, once someone has written it down.

The pricing table is data, not code branches. That choice pays off in the next lesson, when city-specific rules arrive from retrieval with their own percentages.

Examples for the misreads

Two smaller groups also sat at cheap steps. Hinglish misreads (4) and missed veg-meat escalations (2) were cases of the model misreading a kind of input. Step 3 applies: the team added five short worked examples to the contract, three Hinglish tickets with quantities and two veg-meat tickets, each with the correct draft. They also added "chicken", "mutton", "egg", "fish" and "non-veg" to the keyword net, active only when every item in the order is marked vegetarian in the menu data.

Examples are powerful and a little dangerous. The model copies their patterns, including patterns you did not intend. One early example used "2 roti kam" with an order of 4, and the model began assuming 4 rotis in unrelated orders. Keep examples varied, and let the eval set tell you whether they helped.

Check your understanding

0 of 3 answered

1.Six failures come from the model refunding a whole line's price for one missing unit. Which fix addresses the cause?

2.Why is the pricing table written as data (percent and minimum per rule) instead of if-statements?

3.Most of a taxonomy's failures come from missing input data and misplaced arithmetic. What does that suggest about "use the biggest model"?