Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Failure boundaries: when the feature must say "I'm not sure"
In the second week of the pilot, a customer wrote: "found something hard in the pulao, nearly broke my tooth". The assistant drafted a refund of ₹180 for the pulao, with a neat reason. The agent was busy and approved it. Ten days later a lawyer's letter arrived. The refund was not the problem; the missing escalation was. A food-safety report needs the safety team, a call to the restaurant, and a record, not ₹180.
The model did what it was built to do: find the item and compute a refund. Nobody had made it structurally impossible to handle this ticket as a routine refund. A prompt line said "escalate food safety issues", and the model did not see "something hard" as food safety.
A failure boundary is a line that says: beyond here, the feature does not decide. Good boundaries are drawn before launch, enforced in more than one place, and measured.
Three kinds of boundary
The TiffinGo assistant has three kinds.
Out of scope. The ticket is not an order issue at all: a failed payment, an account problem, a complaint about a rider's behaviour. The assistant should not invent a refund; it should route the ticket elsewhere.
Must escalate. The assistant could produce an answer, but policy says a person decides. Food safety (R5), high value (R6), repeat refunds (R7), a vegetarian customer served meat, and any legal or media threat.
Not enough information. "missing items" with no detail, a complaint that does not match the order, or a photo with no text. The honest output is a question for the customer, not a guess.
The output contract carries these as an escalation_code: out_of_scope, food_safety, high_value, repeat_refunds, needs_info or unclear. When the code is set, the action is escalate and refund_inr is 0.
Boundaries that matter belong in code
A prompt instruction is a strong hint, not a guarantee. For the boundaries whose failure is serious or expensive, add checks the model cannot skip.
Two kinds of check are possible. When the boundary depends on structured data, like R6 and R7, code decides exactly. When it depends on language, like R5, code can add a cheap, high-recall safety net: a keyword list that errs towards escalating.
1import re23SAFETY_WORDS = [4 "vomit", "sick", "stomach", "hospital", "doctor", "allerg", "hair", "insect",5 "cockroach", "worm", "plastic", "glass", "metal", "stone", "hard", "tooth",6 "smell", "stale", "rotten", "fungus", "keeda", "baal", "ulti", "kharab",7]8# \b anchors each stem at a word start, so "ulti" does not match "multiple"9SAFETY_RE = re.compile(r"\b(" + "|".join(SAFETY_WORDS) + ")", re.IGNORECASE)1011def forced_escalation(ticket_text: str, order: dict, draft: dict) -> str | None:12 """Return an escalation code the draft must carry, or None."""13 if SAFETY_RE.search(ticket_text):14 return "food_safety"15 if order["order_total_inr"] > 1500 or draft.get("refund_inr", 0) > 500:16 return "high_value"17 if order["refunds_last_30_days"] >= 3:18 return "repeat_refunds"19 return None2021def apply_boundaries(ticket_text: str, order: dict, draft: dict) -> dict:22 code = forced_escalation(ticket_text, order, draft)23 if code and draft.get("escalation_code") != code:24 draft = {**draft, "action": "escalate", "refund_inr": 0, "escalation_code": code,25 "reason": f"System escalation ({code}). Model draft: {draft.get('reason', '')}"}26 return draftThis runs after the model call and before the agent sees the draft. It overrides the model whenever a hard rule applies. The model's original reason is kept, because it helps the person who picks up the ticket.
The keyword list will fire on harmless tickets. "The biryani was hard to find in the bag" contains "hard". That is fine: a false alarm costs a senior agent about four minutes, while a missed food-safety case can cost a lawsuit. The list is a net under the model, not a replacement for it. Words like "baal" (hair), "keeda" (insect), "ulti" (vomit) and "kharab" (spoiled) are there because a third of tickets are Hinglish.
Teach the model that "not sure" is a right answer
Models lean towards answering. If the prompt treats escalation as a last resort, the model will stretch to produce a refund. State the opposite plainly in the contract:
Escalating is a correct answer, not a failure. If the complaint could describeillness, a foreign object or spoiled food, escalate with food_safety even if youare not certain. If you cannot tell which items are affected, escalate withneeds_info and write the one question the agent should ask the customer.The last sentence matters. A bare "needs_info" leaves the agent to start again. A draft that says "Ask which items were missing; the order had 6 items" saves the agent a minute even when it cannot decide.
How much "not sure" is right?
A feature that escalates everything is safe and useless. One that escalates nothing is fast and dangerous. The right rate comes from the costs of the two kinds of mistake.
| Mistake | Example | Cost | Target |
|---|---|---|---|
| Missed escalation | Food-safety ticket drafted as a refund | Legal risk, a harmed customer, lost trust | Close to zero for food safety |
| Unnecessary escalation | "hard to find" sent to the safety team | About 4 minutes of senior time, around ₹25 | Tolerable up to a few per cent |
| Missed "needs info" | Guessing which items were missing | Wrong refund, about ₹80 | Low |
In TiffinGo's labelled data, about 9% of tickets truly need a human: 2% food safety, 3% high value, 3% repeat refunds and about 1% unclear. The team set a target escalation rate of 10 to 14%. A little above 9%, because unnecessary escalations are cheap. Not much higher, because every escalation gives up the time the feature was built to save.
Measure both mistakes separately. A single "accuracy" number hides the one that matters. An assistant that is 92% accurate overall but misses one in four food-safety cases is not ready to launch.
Check your understanding
0 of 3 answered
1.Why does TiffinGo enforce R6 (high value) in code instead of trusting the prompt?
2.The keyword net escalates "the biryani was hard to find in the bag" as food safety. What should the team do?
3.The assistant scores 92% overall on the eval set but misses 2 of 8 food-safety tickets. Is it ready?