Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

When to fine-tune, and when not to


At a quarterly review, a director asked the question every applied-AI team hears sooner or later: "Why are we doing all this retrieval work? Why not fine-tune a model on the 400 policies? Then it would simply know them."

It is a natural idea. Fine-tuning sounds like teaching, and teaching sounds like the way to get knowledge in. But fine-tuning changes a model in a specific way, and adding reliable, citable, up-to-date facts is one of the things it does worst. The team answered the director with numbers, and then fine-tuned something else entirely.

This lesson gives you the decision rule. When fine-tuning is the right tool, it can make a small model as good as a large one at a narrow task, at a fraction of the latency. When it is the wrong tool, it costs weeks and makes things worse.

PolicyPal's three fine-tuning candidatesNo: 38changes a yearNoRetrieval is 0.91RejectedYesNot neededNo complaintsRejectedYes: six routesNot needed835 ms, privacyAcceptedStable?Citable?Named gain?VerdictKnow the policiesHouse styleThe router
Fine-tuning teaches patterns, not facts that change — only the narrow, stable router earned it.

What fine-tuning actually changes

Fine-tuning continues training a model on your examples, adjusting its weights so that your examples become more likely. The model learns patterns from those examples: a format, a tone, a decision rule, a mapping from inputs to labels. That makes it good at a few things.

  • Consistent behaviour and format that prompting struggles to hold across thousands of calls.
  • Narrow tasks such as classification or extraction, where a small model can learn to match a large one.
  • Efficiency: moving a task from a large, slow, remote model to a small, fast, local one.

It is poor at adding facts. The model does absorb some facts from training text, but it blends them with everything else it knows, cannot tell you which document a fact came from, and does not forget the old version when a policy changes. Research on fine-tuning with new knowledge has found that it can make models more likely to state wrong facts confidently. For a system whose core promise is "every sentence has a source", that is disqualifying.

Fine-tuning is good for

  • A fixed output format held across every call
  • A closed-label classifier on a small model
  • A narrow skill that must be fast and cheap
  • Behaviour that is hard to describe but easy to show

Fine-tuning is poor for

  • Facts that change, such as policies and prices
  • Answers that must cite a source
  • Tasks you have not yet measured with a prompt
  • Anything you cannot afford to retrain regularly

Try the cheaper tools first

Most problems that look like "the model needs fine-tuning" are solved earlier on the ladder. Each rung below is cheaper to build and to change than the one after it.

Problem you seeTry firstFine-tune only if
Wrong format or ramblingA clear prompt and a schema (Section 2)Format still drifts at scale with a schema in place
Inconsistent judgement at a boundaryDynamic few-shot examples (Section 2)Hundreds of examples are needed to cover the cases
Does not know your factsRetrieval (Section 3)Never, for facts that change
Needs live data or actionsTools (Section 4)Never; this is what tools are for
Too slow or too costly for a narrow taskA smaller model with a good promptThe small model's measured quality is too low

The last row is where fine-tuning most often earns its place, and it is exactly where PolicyPal's candidate came from.

PolicyPal's three candidates

The team listed every place fine-tuning had been suggested and applied the table.

CandidateWould it help?Verdict
"Know the 400 policies"No. Policies changed 38 times last year; answers must cite sources; retrieval already finds the right section 91% of the time.Rejected
"Answer in Harbourline's style"Barely. The prompt and examples already produce the style; the eval showed no style complaints.Rejected
The routerYes. Six fixed labels, runs on every message, 850 ms of latency today, thousands of logged examples.Accepted

The router stands out on every line. The task is narrow and stable: six routes that change perhaps once a year. It runs before everything, so its 850 milliseconds are added to every single question. The logs from Section 4 already contain thousands of real messages with labels. And there is a benefit that has nothing to do with speed: a local router decides that a message is a sensitive HR topic before any of its text is sent to an external model provider. For a message about a health condition or a harassment complaint, that is a real privacy gain.

The full cost of owning a fine-tuned model

Training is the cheap part. A fair estimate counts everything.

  • Data: about 3,000 labelled messages, of which 820 needed human review, about 25 hours of HR and IT time.
  • Training: about 8 minutes on one cloud GPU, well under a dollar per run, including experiments perhaps $20.
  • Serving: a small GPU instance, or a CPU at higher latency, around $150 to $300 a month, plus deployment and monitoring.
  • Retraining: whenever a route changes or accuracy drifts; plan for a few runs a year, each with a data refresh.
  • Ownership: someone must understand the model, its data and its failure modes after the author moves teams.

Now compare the prompted router: about $0.001 per question, or roughly $26 a month. The fine-tuned router costs more in money. The case for it is not cost. It is 835 milliseconds saved on every question and sensitive text that never leaves the company before routing. If neither of those mattered, the honest answer would be to keep the prompt.

  1. Is the task narrow and stable? — a closed set of outputs that will not change every month.
  2. Have you reached the ceiling of prompting? — measured on a held-out set, not assumed.
  3. Do you have enough good examples? — a few hundred for format, a few thousand for a multi-class classifier.
  4. Can you name the gain in numbers? — milliseconds, dollars, accuracy on a specific class, or data that stays in-house.
  5. Can you own retraining? — data refresh, evaluation and deployment, more than once.

Check your understanding

0 of 3 answered

1.Why is fine-tuning a poor fit for teaching PolicyPal the 400 policies?

2.The fine-tuned router costs more per month than the prompted one. What justifies it?

3.A team's chatbot sometimes rambles past 300 words. What should they try before fine-tuning?