Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Is AI worth it here? The necessity test and the value/risk grid
Two weeks after the demo, TiffinGo's head of product arrived with a list: an assistant that refunds customers automatically, AI replies on WhatsApp, automatic complaint tags for the analytics team, AI-written menu descriptions for restaurants, and "AI delivery times". Every item had an enthusiastic sponsor. The engineering team had capacity for one.
Choosing badly here is expensive in a quiet way. A feature that did not need a model costs more to run and is harder to test than the rules it replaced. A feature that is too risky for its safeguards produces one bad week that makes the company nervous about AI for a year.
This lesson gives you two tools for the choice. The necessity test asks whether a model is needed at all. The value/risk grid asks which of the remaining ideas to build first, and how carefully.
The necessity test
Ask these four questions about any proposed AI feature, in this order. A "no" to any of the first three is a strong sign to use something else.
- Can rules, a form or a lookup do it? If the input is already structured, such as timestamps, prices or a status code, ordinary code is cheaper, faster, and right every time.
- Is the hard part understanding language, images or messy input? This is where models earn their cost. "roti kam thi" defeats any keyword rule; a model reads it easily.
- Can the task tolerate some errors, and is there a way to catch them? Every model feature is wrong sometimes. If one error is a disaster and nothing checks the output, do not ship it.
- Is there enough volume or value to pay for the engineering? An eval set, a validator and monitoring cost weeks. Thirty tickets a day may not justify that.
Here is the test applied to the product head's list.
| Idea | Rules enough? | Messy language? | Errors tolerable and catchable? | Verdict |
|---|---|---|---|---|
| Draft refunds for agents | No | Yes | Yes, an agent reviews each draft | Use a model |
| Refund automatically | No | Yes | Errors reach customers unchecked | Not yet |
| WhatsApp replies | No | Yes | A wrong promise is public | Not yet |
| Complaint tags for analytics | Partly | Yes | Yes, errors average out in charts | Use a model, low priority |
| Delivery-time estimates | Not rules, but not language either | No | Yes | Use a classic ML model on trip data, not an LLM |
The last row matters. "AI delivery times" is a prediction from structured data: distance, time of day, restaurant prep time. That is a regression problem with years of good tools behind it. A large language model would be slower, costlier and worse at it.
Split the feature: code does what code can
The necessity test also works inside one feature. Look at the refund rules again. R4, the late-delivery refund, needs only two timestamps and a subtraction. There is no reason to ask a model to do date arithmetic it might get wrong.
1from datetime import datetime23LATE_LIMIT_MIN = 304LATE_REFUND_INR = 4056def minutes_late(promised_at: str, delivered_at: str) -> int:7 promised = datetime.fromisoformat(promised_at)8 delivered = datetime.fromisoformat(delivered_at)9 return max(0, int((delivered - promised).total_seconds() // 60))1011def late_refund_inr(promised_at: str, delivered_at: str) -> int:12 late = minutes_late(promised_at, delivered_at)13 return LATE_REFUND_INR if late > LATE_LIMIT_MIN else 01415print(late_refund_inr("2026-03-14T20:45:00+05:30", "2026-03-14T20:52:00+05:30")) # 016print(late_refund_inr("2026-03-14T20:45:00+05:30", "2026-03-14T21:31:00+05:30")) # 40The service runs this before calling the model and adds any R4 line to the draft itself. The model still receives minutes_late so its reason can mention it, but it is told not to apply R4. The same thinking applies to R6 and R7: order totals and refund counts are numbers in the database, so code checks them.
What is left for the model is the part only a model can do: reading the complaint, working out which items are affected and how, and spotting when a message is really a food-safety report. That is a smaller, clearer job, and a smaller job is easier to evaluate.
The value/risk grid
For the ideas that pass the necessity test, estimate two numbers.
Value is what the feature saves or earns per day. For drafts: 1,200 tickets × 3.5 minutes saved × ₹250 per agent-hour is about ₹17,500 a day.
Risk is the expected cost of errors that nobody catches:
risk per day = volume x error rate x (1 - catch rate) x cost per errorSuppose drafts are wrong 8% of the time and an uncaught error costs ₹80 on average. With an agent reviewing and catching 85% of errors, the risk is 1,200 × 0.08 × 0.15 × ₹80, about ₹1,150 a day. Automatic refunds have no reviewer, so the catch rate is zero and the risk is 1,200 × 0.08 × 1.0 × ₹80, about ₹7,700 a day, before counting customers who learn to game an automatic system.
High value, low risk
- Build first
- Drafts for agents: about ₹17,500 value, about ₹1,150 risk a day
- Complaint tags: errors blur a chart, nothing more
High value, high risk
- Build only with a human in the loop, or narrow the scope
- Automatic refunds: about ₹7,700 risk a day, plus fraud
- WhatsApp replies: a wrong promise is public and permanent
Low-value, low-risk ideas such as menu descriptions are fine experiments when the team has spare time. Low-value, high-risk ideas should simply be dropped.
The numbers are rough, and that is fine. Their job is to make the conversation concrete. "Automatic refunds are risky" starts an argument. "Automatic refunds carry about ₹7,700 a day of unreviewed error, six times the draft version" ends one.
When the answer is no
Be ready to say no, and to say why. Do not use a model when rules cover the input, as with late deliveries. Do not use one when an error is severe and no one checks the output. Do not use one when volume is so low that a person is cheaper than the engineering, or when you have no real examples to evaluate against, because then you cannot know whether it works.
"Not yet" is also an answer. Automatic refunds may become sensible for one narrow case, such as a single missing item under ₹100 for customers with no recent refunds, once months of draft data show the assistant is right on that case 99% of the time.
Check your understanding
0 of 3 answered
1.Which part of the refund decision should TiffinGo's service compute without the model?
2.Drafts are wrong 8% of the time, agents catch 85% of errors, and an uncaught error costs ₹80. What is the daily risk at 1,200 tickets?
3.The "AI delivery times" idea fails the necessity test mainly because: