AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Is AI worth it here? The necessity test and the value/risk grid


Two weeks after the demo, TiffinGo's head of product arrived with a list: an assistant that refunds customers automatically, AI replies on WhatsApp, automatic complaint tags for the analytics team, AI-written menu descriptions for restaurants, and "AI delivery times". Every item had an enthusiastic sponsor. The engineering team had capacity for one.

Choosing badly here is expensive in a quiet way. A feature that did not need a model costs more to run and is harder to test than the rules it replaced. A feature that is too risky for its safeguards produces one bad week that makes the company nervous about AI for a year.

This lesson gives you two tools for the choice. The necessity test asks whether a model is needed at all. The value/risk grid asks which of the remaining ideas to build first, and how carefully.

Where TiffinGo's five ideas landedAgent drafts — buildAutorefunds — not yetMenu text —experimentDrop itLow riskHigh riskHigh valueLow valueDaily uncaught-error risk: drafts about ₹1,150, auto refunds about ₹7,700.
The same model is low risk as a reviewed draft and high risk as an automatic refund — the grid scores the product, not the model.

The necessity test

Ask these four questions about any proposed AI feature, in this order. A "no" to any of the first three is a strong sign to use something else.

  1. Can rules, a form or a lookup do it? If the input is already structured, such as timestamps, prices or a status code, ordinary code is cheaper, faster, and right every time.
  2. Is the hard part understanding language, images or messy input? This is where models earn their cost. "roti kam thi" defeats any keyword rule; a model reads it easily.
  3. Can the task tolerate some errors, and is there a way to catch them? Every model feature is wrong sometimes. If one error is a disaster and nothing checks the output, do not ship it.
  4. Is there enough volume or value to pay for the engineering? An eval set, a validator and monitoring cost weeks. Thirty tickets a day may not justify that.

Here is the test applied to the product head's list.

IdeaRules enough?Messy language?Errors tolerable and catchable?Verdict
Draft refunds for agentsNoYesYes, an agent reviews each draftUse a model
Refund automaticallyNoYesErrors reach customers uncheckedNot yet
WhatsApp repliesNoYesA wrong promise is publicNot yet
Complaint tags for analyticsPartlyYesYes, errors average out in chartsUse a model, low priority
Delivery-time estimatesNot rules, but not language eitherNoYesUse a classic ML model on trip data, not an LLM

The last row matters. "AI delivery times" is a prediction from structured data: distance, time of day, restaurant prep time. That is a regression problem with years of good tools behind it. A large language model would be slower, costlier and worse at it.

Split the feature: code does what code can

The necessity test also works inside one feature. Look at the refund rules again. R4, the late-delivery refund, needs only two timestamps and a subtraction. There is no reason to ask a model to do date arithmetic it might get wrong.

Python
from datetime import datetimeLATE_LIMIT_MIN = 30LATE_REFUND_INR = 40def minutes_late(promised_at: str, delivered_at: str) -> int:    promised = datetime.fromisoformat(promised_at)    delivered = datetime.fromisoformat(delivered_at)    return max(0, int((delivered - promised).total_seconds() // 60))def late_refund_inr(promised_at: str, delivered_at: str) -> int:    late = minutes_late(promised_at, delivered_at)    return LATE_REFUND_INR if late > LATE_LIMIT_MIN else 0print(late_refund_inr("2026-03-14T20:45:00+05:30", "2026-03-14T20:52:00+05:30"))  # 0print(late_refund_inr("2026-03-14T20:45:00+05:30", "2026-03-14T21:31:00+05:30"))  # 40

The service runs this before calling the model and adds any R4 line to the draft itself. The model still receives minutes_late so its reason can mention it, but it is told not to apply R4. The same thinking applies to R6 and R7: order totals and refund counts are numbers in the database, so code checks them.

What is left for the model is the part only a model can do: reading the complaint, working out which items are affected and how, and spotting when a message is really a food-safety report. That is a smaller, clearer job, and a smaller job is easier to evaluate.

The value/risk grid

For the ideas that pass the necessity test, estimate two numbers.

Value is what the feature saves or earns per day. For drafts: 1,200 tickets × 3.5 minutes saved × ₹250 per agent-hour is about ₹17,500 a day.

Risk is the expected cost of errors that nobody catches:

Text
risk per day = volume x error rate x (1 - catch rate) x cost per error

Suppose drafts are wrong 8% of the time and an uncaught error costs ₹80 on average. With an agent reviewing and catching 85% of errors, the risk is 1,200 × 0.08 × 0.15 × ₹80, about ₹1,150 a day. Automatic refunds have no reviewer, so the catch rate is zero and the risk is 1,200 × 0.08 × 1.0 × ₹80, about ₹7,700 a day, before counting customers who learn to game an automatic system.

High value, low risk

  • Build first
  • Drafts for agents: about ₹17,500 value, about ₹1,150 risk a day
  • Complaint tags: errors blur a chart, nothing more

High value, high risk

  • Build only with a human in the loop, or narrow the scope
  • Automatic refunds: about ₹7,700 risk a day, plus fraud
  • WhatsApp replies: a wrong promise is public and permanent

Low-value, low-risk ideas such as menu descriptions are fine experiments when the team has spare time. Low-value, high-risk ideas should simply be dropped.

The numbers are rough, and that is fine. Their job is to make the conversation concrete. "Automatic refunds are risky" starts an argument. "Automatic refunds carry about ₹7,700 a day of unreviewed error, six times the draft version" ends one.

When the answer is no

Be ready to say no, and to say why. Do not use a model when rules cover the input, as with late deliveries. Do not use one when an error is severe and no one checks the output. Do not use one when volume is so low that a person is cheaper than the engineering, or when you have no real examples to evaluate against, because then you cannot know whether it works.

"Not yet" is also an answer. Automatic refunds may become sensible for one narrow case, such as a single missing item under ₹100 for customers with no recent refunds, once months of draft data show the assistant is right on that case 99% of the time.

Check your understanding

0 of 3 answered

1.Which part of the refund decision should TiffinGo's service compute without the model?

2.Drafts are wrong 8% of the time, agents catch 85% of errors, and an uncaught error costs ₹80. What is the daily risk at 1,200 tickets?

3.The "AI delivery times" idea fails the necessity test mainly because: