AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

What you will build: the TiffinGo Order Issue Assistant


It is 8:58 pm on a Friday in Bengaluru. The TiffinGo support queue has 212 open tickets, and Priya, a support agent, opens the next one: "dal was cold, one roti missing. ordered 4 only 3 came. very bad". To answer it she opens the order, finds what the customer paid for each item, looks up the refund policy, works out an amount, and types a reason for the audit log. That takes her about six minutes.

With the assistant you will build in this course, the same ticket opens with a draft already waiting: refund ₹96, one line per item, and the policy rule behind each line. Priya reads it, checks it against the order, and clicks Approve. That takes about seventy seconds. The money only moves when she clicks.

This lesson shows you the finished product before we build any of it. Every later lesson adds one piece to this system, so it helps to know where we are going.

One complaint, from ticket to approved refundTickettext plusorder recordPrompt contractdrafts items and rulesValidatorchecks itagainst the orderCode priceslines,applies capsAgent approves₹96 in about 70 sSix minutes by hand; about seventy seconds with a draft.
The model is one box in the middle; money moves only at the last box, after a person clicks.

The problem, in numbers

TiffinGo is a food-delivery app in six Indian cities. It delivers about 40,000 orders a day, and about 3% of orders produce a complaint. That is 1,200 order-issue tickets a day, rising to about 2,000 on a festival weekend.

Support has 45 agents across shifts. A ticket takes six minutes on average, so the team spends 1,200 × 6 = 7,200 minutes, or 120 agent-hours a day, on these tickets. At a loaded cost of about ₹250 per agent-hour, that is ₹30,000 a day.

Time is only half the problem. An audit of 200 "cold food" tickets found refunds from ₹0 to the full order value for the same complaint. Two agents read the same policy and made different decisions. Customers noticed, and some learned which words got the biggest refund.

So the product goal has two parts. Cut handling time from six minutes to about two and a half. And make refunds follow the policy the same way every time.

The refund policy the assistant applies

TiffinGo's refund policy is long, but its core fits in seven rules. We will refer to them by number for the whole course.

RuleSituationWhat to do
R1Item missingRefund what the customer paid for the missing units
R2Food cold or badly packedRefund 30% of the affected item's paid price, minimum ₹30
R3Wrong item deliveredReplace if the restaurant is open, otherwise refund the item
R4Delivered more than 30 minutes after the promised timeFlat ₹40 refund
R5Food safety: foreign object, spoiled food, illnessAlways escalate to a human
R6Draft refund above ₹500, or order above ₹1,500Escalate for approval
R7Customer has 3 or more refunds in the last 30 daysEscalate for review

The first four rules produce an amount. The last three say when the assistant must stop and hand the ticket to a person. That split, between work the model may draft and work it must never decide, is a theme you will see again and again.

What the finished assistant does

The assistant receives two things: the customer's words, and the order record from TiffinGo's order service.

JSON
{  "ticket_text": "dal was cold, one roti missing. ordered 4 only 3 came. very bad",  "order": {    "order_id": "TG-48213",    "city": "Bengaluru",    "promised_at": "2026-03-14T20:45:00+05:30",    "delivered_at": "2026-03-14T20:52:00+05:30",    "items": [      {"name": "Dal Makhani", "qty": 1, "paid_inr": 220},      {"name": "Butter Roti", "qty": 4, "paid_inr": 120},      {"name": "Jeera Rice", "qty": 1, "paid_inr": 150}    ],    "order_total_inr": 490,    "refunds_last_30_days": 0  }}

It returns a draft in a fixed shape that code can check, with one line per affected item:

JSON
{  "action": "refund",  "lines": [    {"item": "Butter Roti", "units": 1, "issue": "missing", "rule": "R1", "amount_inr": 30},    {"item": "Dal Makhani", "units": 1, "issue": "cold", "rule": "R2", "amount_inr": 66}  ],  "refund_inr": 96,  "reason": "1 of 4 Butter Roti missing: ₹30 (R1). Dal Makhani cold: 30% of ₹220 = ₹66 (R2).",  "escalation_code": null,  "question_for_customer": null}

Check the arithmetic yourself. Four rotis cost ₹120, so one missing roti is ₹30 under R1. Thirty per cent of ₹220 for the cold dal is ₹66 under R2. The total is ₹96. The order arrived seven minutes after the promised time, which is under the 30-minute limit in R4, so there is no late refund.

The agent then sees this screen:

Text
TICKET 88121  ·  TG-48213  ·  Bengaluru  ·  8:58 pm"dal was cold, one roti missing. ordered 4 only 3 came. very bad"SUGGESTED: Refund Rs 96                    prompt v3 · policy 2026-03  Butter Roti   1 of 4 missing     Rs 30   (R1 missing item)  Dal Makhani   cold               Rs 66   (R2 cold food, 30%)  Not late: delivered 7 min after the promised time.[ Approve Rs 96 ]   [ Edit ]   [ Escalate ]   [ Reject draft ]

Notice what the screen does not do. It does not message the customer. It does not move money on its own. It shows its working, so Priya can check a draft in seconds instead of redoing the work.

What sits around the model call

The call to the model is perhaps forty lines of code. Everything that makes it trustworthy is around it, and that is what this course is about.

PartWhat it doesBuilt in
Scope and failure boundariesDecides what the assistant may draft and when it must escalateSection 2
Prompt contract and validatorFixes the input, the output shape and the rules; rejects bad draftsSection 3
Eval set of 120 real ticketsMeasures quality before every changeSection 4
Policy retrieval and order-status toolGives the model facts it lacksSection 5
Monitoring, fallbacks, launch reviewKeeps it working after launchSection 6

The course map

  1. Scope — decide whether AI is worth it here, and where the feature must say "I'm not sure".
  2. Interface — write the prompt as a contract, get structured output, validate it, and version it.
  3. Evaluate — build the 120-ticket eval set, grade outputs, analyse failures, and gate every release.
  4. Extend — add retrieval, a tool, or a different model only when a measured failure calls for it.
  5. Operate — defend against misuse, survive outages, watch drift, cost and latency, and run a launch review.

Each stage is one section of the course, and each section leaves the assistant a little more complete. By the end, you will have built every part of the system described above, and you will know why each part is there.

Check your understanding

0 of 3 answered

1.Why does the assistant produce a draft for an agent instead of issuing refunds directly?

2.The order arrived at 8:52 pm against a promised 8:45 pm. Which rule does this trigger?

3.What share of the assistant's code is the model call itself?