Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Fine-tuning: when it helps, and when it is a trap
A vendor presentation ended with a clear pitch: "You have 50,000 historical tickets with agent decisions. Fine-tune on them, and the model will learn your policy. No more long prompts." It sounded efficient. It would also have taught the assistant exactly the wrong thing.
Remember the audit in the first lesson: for the same "cold food" complaint, agents had refunded anything from ₹0 to the full order value. Fifty thousand historical decisions are fifty thousand examples of that inconsistency. A model trained on them would learn to be inconsistent in the same way, and would do it with the confidence of a model rather than the hesitation of a new agent.
Fine-tuning is a real tool with real uses. It is also the most expensive capability on the ladder, and the one most often reached for too early.
What fine-tuning changes
Fine-tuning is good at changing behaviour: the format of the output, the tone, where the boundary sits between two categories, what to do with a style of input it sees often. It is poor at storing facts that must be exact and must change on a known date. A fine-tuned model may remember that Mumbai has a monsoon rule; it cannot be told that the rule ends on 30 September except by training again.
It also removes a lever you rely on. With a prompted model, a policy change is a new prompt release, evaluated and rolled out in days. With a fine-tuned model, it is a new training run, new evaluation and a new model to host, and until it finishes the model keeps applying the old policy.
Why it is a trap for TiffinGo today
Run the TiffinGo situation through five questions.
| Question | TiffinGo's answer | Points towards |
|---|---|---|
| Is the task stable? | Policy changes twice a month; city rules come and go | Prompt and retrieval |
| Is the training data clean? | Historical decisions are inconsistent | Not fine-tuning |
| Is the prompted model failing on behaviour that examples cannot fix? | No; remaining failures are live data and edge cases | Not fine-tuning |
| Is cost or latency a real problem? | About ₹745 a day for the model; latency hidden by pre-generating drafts | Not yet |
| Is volume large enough to repay the effort? | 1,200 tickets a day | Not yet |
Every row says "not now". That is typical for a feature in its first year, and it is why fine-tuning comes last on the ladder.
When it does help
Fine-tuning becomes attractive when the task is stable, you have clean examples, and a smaller or cheaper model would be good enough if it learned your pattern. The common, sensible case is distillation: a strong prompted model produces outputs, humans approve them, and those approved outputs train a smaller model to do the same narrow job faster and cheaper.
It also helps with narrow classification at high volume, such as tagging every one of 40,000 daily orders' reviews by topic, where a small fine-tuned model is far cheaper than a large prompted one and the label set rarely changes.
Clean data comes from production
If TiffinGo does fine-tune one day, the training data will not be the 50,000 historical tickets. It will be the assistant's own drafts that agents approved unchanged under the current policy, the cleanest labels the company has ever had.
1import json23def training_rows(decisions, eval_ids: set, policy_version: str):4 """Yield chat-format training examples from logged, human-approved drafts."""5 for d in decisions:6 if d["agent_outcome"] != "approved_unchanged":7 continue # only drafts a human accepted as they were8 if d["policy_version"] != policy_version:9 continue # never teach a policy that is no longer in force10 if d["ticket_id"] in eval_ids:11 continue # never train on the tickets you evaluate with12 yield {"messages": [13 {"role": "system", "content": d["system_prompt"]},14 {"role": "user", "content": d["model_input"]},15 {"role": "assistant", "content": json.dumps(d["model_output"], ensure_ascii=False)},16 ]}Each filter removes a specific poison. Edited drafts contain the model's mistakes. Old-policy drafts teach rules that no longer apply. And eval tickets in the training data make the eval score meaningless, because the model has seen the answers. The chat-message format shown is the common shape most fine-tuning services accept; check your provider's exact format.
Balance matters too. Approved drafts are mostly routine missing-item refunds. If escalations are 10% of the data, the fine-tuned model may learn to escalate less. Check the category mix against the eval set's table before training.
Compare fairly, on the same gate
A fine-tuned model is just another candidate release. It goes through the same eval set, the same critical-category rule and the same flip report as any prompt change. Compare it against the prompted baseline on quality, cost and latency together, not on the one number the vendor quotes.
Check your understanding
0 of 3 answered
1.Why would fine-tuning on 50,000 historical tickets harm TiffinGo's assistant?
2.Why does training_rows skip tickets that are in the eval set?
3.When is fine-tuning most likely to pay off?