Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
The lifecycle: scope, interface, evaluate, extend, operate
When the TiffinGo team saw 58% on real tickets, the first suggestion in the meeting was "let's try a bigger model". The second was "let's add RAG". Both were guesses. Nobody yet knew why the 42% failed, so nobody could know whether a bigger model or retrieval would help. They might have spent a month and a lakh of rupees on the wrong fix.
A lifecycle is a fixed order of questions. It stops a team from reaching for the most exciting fix before the boring questions have answers. This course uses five stages, and each section of the course is one stage.
The stages are not new ideas. What is specific to LLM features is that each stage leaves behind an artefact, a file or document you keep, and the next stage depends on it.
The five stages
- Scope — decide whether AI is the right tool, what the feature may decide, and where it must stop. Artefact: a one-page scope with the value/risk position and the escalation rules.
- Interface — define exactly what goes into the model and what must come out. Artefact: a versioned prompt contract, an output schema and a validator.
- Evaluate — measure quality on a fixed set of real cases. Artefact: the eval set, the graders, a failure taxonomy and a regression gate.
- Extend — add a capability (retrieval, tools, a different model, fine-tuning) only to fix a measured failure. Artefact: a decision record linking each capability to the failures it fixed.
- Operate — run it safely at volume. Artefact: dashboards, alerts, a fallback plan and a launch checklist.
Each artefact is small. The scope is one page. The first prompt contract is about sixty lines. The eval set starts at 120 rows. But without them the next stage has nothing to stand on.
Why the order matters
The order is the point. You cannot evaluate before you have an interface, because there is no fixed output to grade. You should not extend before you evaluate, because you do not know what is broken. And you should not scope after you build, because by then the scope is whatever the code happens to do.
Look at the two suggestions from the meeting through this lens. "Try a bigger model" and "add RAG" are both Extend-stage moves. The team was at the end of the Interface stage with no eval set. The right next step was Evaluate: build the set, grade it, and read the failures.
When TiffinGo did that, the largest failure group was wrong prices: the model used menu prices instead of what the customer actually paid after a discount. That is an Interface problem, fixed by putting paid_inr into the input. It cost one afternoon. A bigger model would have made the same mistake, because the right number was not in the prompt.
It is a loop, not a line
The stages form a loop. Production sends you back to earlier stages, and that is normal.
| Signal in production | Stage it sends you back to | TiffinGo example |
|---|---|---|
| A new kind of complaint appears | Scope | Customers start reporting "cutlery missing"; is that in scope? |
| Drafts fail validation more often | Interface | A menu change adds items with quantities like "half plate" |
| Agents edit more drafts than last month | Evaluate | Add the edited tickets to the eval set and re-run |
| Many failures need facts the model lacks | Extend | A new monsoon policy for Mumbai deliveries |
| Cost per ticket doubles | Operate, then Extend | Prompt grew; review model choice or trim context |
A healthy team goes around the loop many times. Each time, the eval set grows a little and the scope gets sharper.
Using the lifecycle to place a problem
The most useful habit this lifecycle builds is asking "which stage does this problem belong to?" before choosing a fix. Here are four problems from TiffinGo's first month, placed correctly.
- "The assistant refunded a customer who found a hair in their biryani." Scope. Food safety should never have been something the assistant could refund. The fix is a failure boundary, enforced in code, not a better prompt.
- "Sometimes the reply is not valid JSON." Interface. Use structured output and a validator with a clear error path.
- "The new prompt seems better." Evaluate. "Seems" is not a measurement. Run the eval set and compare per category.
- "It gets Mumbai rain-damage refunds wrong." Extend, probably. The Mumbai monsoon rule is not in the prompt, so retrieval over the policy may fix it. Confirm with the eval set first.
Placing the problem first often makes the fix smaller. A food-safety refund looks like a model mistake, but it is really a missing boundary.
Check your understanding
0 of 3 answered
1.The team has a working prompt but no eval set. Someone proposes switching to a larger model. What should happen first?
2.The assistant drafted a refund for a customer who reported food poisoning. Which stage owns the fix?
3.Why does each stage produce an artefact, such as an eval set or a prompt contract?