Fine-Tuning LLMs with LoRA, QLoRA and PEFT

When and Why to Fine-Tune — and When Not To


A fintech team spent six weeks and about 40,000 dollars fine-tuning a 13B model to classify support tickets into 14 categories. Accuracy came out at 91%. A week after launch, a sceptical engineer spent an afternoon writing a careful prompt for the off-the-shelf model — clear category definitions, six worked examples, an explicit "if unsure, output needs_review" instruction. It scored 89%.

Two points of accuracy for six weeks and 40,000 dollars. And the prompt could be edited in thirty seconds when Legal added a fifteenth category, whereas the fine-tune needed a fresh labelling round and a retraining run.

That is not an argument against fine-tuning. It is an argument for answering one question honestly before you start: what specifically is the base model failing at, and is fine-tuning the cheapest thing that fixes it? Most teams skip that question, and most wasted fine-tuning budgets trace back to skipping it.

Name the failure before you pick the fixFine-tuning fixes this• A format the model will not hold• Domain vocabulary itconsistently misreads• Drift across repeated identical calls• A 7B doing a frontier model's narrow jobFine-tuning cannot fix this• Facts that aremissing — that is retrieval• Under a few hundred clean examples• Requirements that change every week• Per-user behaviourchosen at request time
Every failure on the right survives training because it is a knowledge or freshness problem wearing a behaviour problem's clothes.

Start by naming the failure

Every fine-tuning project should begin with a written sentence of this shape: "The base model, given our best prompt, fails on X in Y% of cases, and the business cost of that is Z." If you cannot write that sentence with real numbers in it, you do not yet know whether fine-tuning will help.

Run this before anything else:

  1. Build an evaluation set of 200-500 real examples with correct answers. This is unavoidable work — you need it to measure the fine-tune anyway.
  2. Score the base model with a naive prompt.
  3. Score it with a carefully engineered prompt including 5-10 examples.
  4. Score it with retrieval attached, if the failures look knowledge-shaped.

Now you have a baseline curve. If naive scores 62%, engineered prompt scores 88%, and your target is 90%, you are two points away and fine-tuning is a plausible last mile. If engineered scores 63%, prompting has hit a wall and fine-tuning is the right tool. If engineered scores 94%, you are done — ship it.

Fine-tuning without a prompt-engineered baseline is not engineering, it is spending. You cannot claim an improvement you never measured a starting point for.

The failure modes that fine-tuning actually fixes

Output format and structure the model will not hold

You need strict JSON with a fixed schema, or a specific report layout, on every single call. A good prompt gets you to 96% schema compliance. Fine-tuning on 1,500 examples gets you to 99.8%. When 4% of calls failing means 4,000 broken downstream jobs a day, that difference is the whole project. Format compliance is the single most reliable thing fine-tuning buys.

Domain vocabulary and conventions the base model misreads

In a clinical note, "the patient is negative for chest pain" is a normal finding, not a bad outcome. In shipping logistics, "the container is clean" refers to paperwork, not cleanliness. Base models trained on general web text systematically misread this kind of in-domain convention, and no amount of prompt instruction reliably fixes hundreds of such conventions. Fine-tuning on real in-domain text does.

Consistency across repeated calls

Ask a base model the same borderline question ten times at temperature 0.7 and you may get six "approve" and four "escalate". A model fine-tuned on consistently-labelled examples collapses that variance sharply. If your process requires that identical inputs produce identical decisions — regulated lending, medical triage, content moderation appeals — this is a real requirement, not a nice-to-have.

Latency and token cost at volume

A 12-shot prompt costs the same tokens on every request, forever. Fine-tuning bakes the examples into the weights so the prompt shrinks. Work the numbers:

Few-shot promptingFine-tuned model
Instructions + examples1,800 tokens150 tokens
User query200 tokens200 tokens
Input tokens per request2,000350
At 1M requests/month2,000M tokens350M tokens
At 0.60 dollars per M input tokens1,200 dollars/month210 dollars/month
Time-to-first-token impactFull prefill of 2,000 tokensPrefill of 350 tokens

That is 990 dollars a month saved, or 11,880 a year. Before you bank it, check whether your provider caches a repeated prompt prefix: most major APIs now do, billing cached input tokens at a steep discount and skipping most of their prefill. A long, fixed few-shot prefix is exactly what caching is for, so price the prompting column with caching turned on, or the comparison flatters fine-tuning. Set against it: 2,000 labelled examples at 0.40 dollars each is 800 dollars, GPU time for a LoRA run is roughly 4 hours on an A100 at 1.80 dollars an hour — about 7 dollars — and three weeks of an engineer's time, which at a loaded cost of 5,000 dollars a week is 15,000 dollars. Total roughly 15,800 dollars against 11,880 a year of savings: a 16-month payback on cost alone. At 10M requests a month the payback is under two months and the decision is obvious. At 10,000 requests a month it never pays back.

Small models doing a big model's job

A fine-tuned 7B model frequently matches a general 70B model on one narrow task, at roughly a tenth of the serving cost and a fraction of the latency. This is one of the strongest cases for fine-tuning and the one teams most often overlook: the goal is not "better than the big model at everything", it is "as good as the big model at the one thing we actually do".

The failure modes fine-tuning does not fix

Missing or changing facts

This is the mistake that costs the most. The model does not know your Q3 pricing, so someone fine-tunes it on the price list. Two things go wrong. First, facts absorbed through fine-tuning are unreliable — the model interpolates between them and produces confident, plausible, wrong numbers. Second, prices change in October, and now you need another training run. Retrieval solves this properly: put the price list in a vector store, fetch the relevant rows, put them in the context. Facts belong in context; behaviour belongs in weights.

Fine-tune to change how the model behaves. Retrieve to change what the model knows. Confusing the two produces a model that is confidently wrong and expensive to correct.

Very small datasets

With 80 examples, a LoRA run will drive training loss towards zero and learn essentially nothing generalisable. You will see near-perfect training metrics and no improvement on held-out data. Below roughly 500 clean examples, few-shot prompting almost always wins; 500-1,000 is the grey zone where a low-rank adapter with strong regularisation sometimes helps; above a few thousand, fine-tuning starts to be clearly better.

Requirements that change weekly

If your category taxonomy, tone guidelines or policy rules are still moving, every change invalidates your training data. A prompt is a text file you edit; a fine-tune is a pipeline you re-run. Wait until the specification stabilises.

Per-user personalisation at request time

You cannot fine-tune per user per session. That is what context is for — pass the user's preferences and history in the prompt.

A model that is already good enough

If prompting gets you to 94% and the target is 90%, additional accuracy has no business value. Stop.

Fine-tuning against the alternatives

DimensionPrompt engineeringFew-shot in contextRAGFine-tuning
Time to first resultMinutesMinutesDaysWeeks
Upfront costNear zeroNear zeroModerate (index + pipeline)High (data + compute + time)
Per-request token costLowHigh (examples every call)High (retrieved chunks)Lowest
Handles changing factsManual editManual editExcellent — update the indexPoor — needs retraining
Enforces output formatGoodGoodNo effectExcellent
Learns domain style/toneLimitedLimitedNo effectExcellent
Consistency across callsModerateModerateModerateHigh
Explainability of answersLowLowHigh — cite the sourceLow
ReversibilityInstantInstantInstantRedeploy previous weights

The column headings are not mutually exclusive. The strongest production systems combine them: a fine-tuned model that reliably emits the right structure and speaks the domain's language, fed with retrieved facts so its knowledge is current and citable. Fine-tuning handles the how; retrieval handles the what.

Few-shot versus fine-tuning specifically

These two are the closest substitutes, because both teach by example. The differences that matter:

  • Number of examples the model can see. Few-shot fits maybe 20-50 examples in context before cost and attention dilution bite. Fine-tuning digests 10,000 without changing the prompt at all.
  • Where the cost lands. Few-shot pays per request, forever. Fine-tuning pays once, upfront.
  • Order sensitivity. Few-shot output can shift measurably when you reorder the examples — a real, documented instability. Fine-tuned weights have no such sensitivity.
  • Iteration speed. Few-shot: seconds. Fine-tuning: hours to days.

Three decisions, worked

Customer support automation

Task: route incoming tickets into 14 categories and draft a first reply. Volume 1.2M tickets a month. Available data: 4 years of resolved tickets with human-written replies — around 800,000 examples. Base model with an engineered prompt reaches 87% routing accuracy but the drafted replies do not sound like the company.

Decision: fine-tune, plus retrieval. The data volume is far past the threshold, the task is stable, the volume makes token savings material, and "sounds like our company" is exactly a style problem — the one thing fine-tuning is best at. Attach retrieval over the current help-centre articles so the replies cite live policy rather than four-year-old policy baked into weights.

Medical research assistant

Task: answer clinician questions about drug interactions from a literature corpus that grows weekly. Roughly 3,000 queries a month. Errors are dangerous, and answers must be traceable to a source.

Decision: RAG, not fine-tuning. The knowledge changes weekly, which rules out baking it into weights. Traceability is a hard requirement and only retrieval provides it. Volume is far too low for token savings to matter. A light fine-tune later, purely to enforce answer structure and hedging language, would be reasonable — but the substance must come from retrieval.

Real-time product search reranking

Task: rerank 50 candidate products against a query, with a 60 ms latency budget. 8M queries a day. The catalogue changes hourly.

Decision: fine-tune a small model. Latency rules out a large model with a long prompt outright. At 8M queries a day the token economics are overwhelming. And note that the changing catalogue is not a problem here, because the products come in as input; what is being learned is a ranking behaviour, which is stable. This is the ideal fine-tuning shape: high volume, tight latency, stable behaviour, dynamic inputs.

Making the call in practice

Work through these in order, and stop at the first one that answers your question:

CheckIf yes
Does an engineered prompt already hit the target?Ship it. Do not fine-tune.
Are the failures about missing or stale facts?Build retrieval first.
Do you have fewer than ~500 clean examples?Few-shot, and start collecting data.
Is the task specification still changing weekly?Wait. Prompt in the meantime.
Are failures about format, tone, domain convention or consistency?Fine-tune. This is the sweet spot.
Is volume high enough that a shorter prompt pays for the project?Fine-tune, and compute the payback month explicitly.
Do you need a small model to match a large one on one task?Fine-tune the small model.

The most common way this goes wrong is treating fine-tuning as a maturity signal rather than a tool — the team wants to have fine-tuned a model, so the analysis is run backwards from that conclusion. The second most common is fine-tuning to fix hallucination. It does not; it usually makes hallucination more confident and more fluent, because the model learns your answer style without acquiring the facts to fill it with.

Fine-tuning changes how a model speaks, not how much it knows. If the complaint is "it makes things up", the fix lives in retrieval and in evaluation, not in the training loop.

What this means when you plan a project

Budget the evaluation set before the training run. Concretely: 300 labelled examples that nothing trains on, scored against your best prompt, committed to version control, and re-scored after every change. That artefact is what turns "the fine-tune seems better" into a number, and it is the thing that lets you cancel the project cheaply in week one instead of expensively in week six.

Then write down the payback calculation with your actual request volume, your actual labelling cost and your actual engineering weeks. Do it before you start, not after. If the payback is longer than the time you expect the requirements to stay stable, the project is not worth doing regardless of how good the model would be — because you will be retraining it before it has paid for itself.