Fine-Tuning LLMs with LoRA, QLoRA and PEFT

Full Fine-Tuning vs. Parameter-Efficient Fine-Tuning (PEFT)


A team decides to adapt Llama-2-7B to write radiology report summaries. They have 4,000 labelled examples and one cloud GPU with 40 GB of memory. They write the obvious training loop, call model.train(), and thirty seconds later CUDA reports out of memory while allocating a 28 GB tensor. They rent a bigger GPU. Same error. They rent two 80 GB A100s and shard the model across them. It runs, barely, at 11 seconds per step.

Then the product manager asks for a second model — same base, but for cardiology. And a third for pathology. Each finished model is a 13.5 GB file. Thirty specialities means 405 GB of near-identical weights sitting in object storage, and thirty separate GPU deployments because you cannot hold thirty 7B models in memory at once.

Nothing about that story is a bug. It is what happens when you update every parameter in a large language model. The entire field of parameter-efficient fine-tuning exists because someone looked at that 405 GB and asked whether all of it was really carrying new information. The answer, it turns out, is no — and the gap between "no" and "how much less" is enormous.

Where the 40 GB goes on a 7B modelFull fine-tuning• Weights 13.5 GB, all of them trainable• Gradients another 13.5 GB in BF16• Adam moments 54 GB in FP32• One 27 GB checkpoint per taskLoRA, rank 16• Same 13.5 GB of weights, frozen• Gradients for 4.2 M params — 8 MB• Adam moments about 50 MB• One 16 MB adapter per task
The base weights are the same size either way; what PEFT removes is the optimiser state stacked on top of them.

Full fine-tuning: what actually gets stored

Full fine-tuning means taking every weight in a pretrained model and making it trainable. You feed in your task data, compute the loss, backpropagate, and let the optimiser nudge all of them. Conceptually it is the same loop that pretrained the model in the first place, just with a smaller dataset and a smaller learning rate.

The cost is not in the maths. It is in what has to live in GPU memory simultaneously. Take Llama-2-7B, which has 6,738,415,616 parameters — call it 6.74 billion. (It is an older model, but its plain shapes make the arithmetic easy to follow, so this course uses it as the worked example; the same method applies to any current model.) Train it with AdamW in the standard mixed-precision setup and here is the bill, in bytes per parameter:

What is storedPrecisionBytes/paramFor 6.74B params
Model weights (compute copy)BF16213.5 GB
GradientsBF16213.5 GB
AdamW first moment (momentum)FP32427.0 GB
AdamW second moment (variance)FP32427.0 GB
FP32 master copy of weightsFP32427.0 GB
Total, before activations—16108 GB

Sixteen bytes per parameter is the number to memorise. Multiply by parameter count and you get the floor of your memory budget before a single activation is stored. Activations for a batch of 4 sequences of 1,024 tokens add another 10-20 GB on top, so the practical requirement is closer to 125 GB — which is why the team above could not fit it on 80 GB and needed two cards.

Full fine-tuning costs roughly 16 bytes of GPU memory per parameter with AdamW in mixed precision. The model weights are the smallest line item on the invoice.

The reason the optimiser dominates is that AdamW keeps two running statistics per parameter, both in FP32 for numerical stability, plus a full-precision master copy of the weights so that tiny updates are not lost to BF16 rounding. Three FP32 tensors the size of the model. That is 12 of the 16 bytes.

What full fine-tuning buys you

  • Maximum capacity to change behaviour. Every weight can move, so the model can learn genuinely new representations — a new language, a new modality of reasoning, a tokeniser-level shift in what the text even looks like.
  • No architectural constraint. You are not assuming the update has any particular structure.
  • Best achievable quality on very large, very different datasets. If you have 500 million tokens of legal Hindi and want a model that thinks in it, full fine-tuning wins.

What it costs you

  • Memory, as computed above.
  • Storage per task. One 13.5 GB artefact per fine-tune, forever.
  • Catastrophic forgetting. With all weights free to move and a small dataset, the model overwrites general capability. A model fine-tuned on 4,000 radiology summaries often gets measurably worse at arithmetic, at following instructions, at everything it was not shown.
  • Sensitivity. A learning rate that is 3× too high on full fine-tuning destroys the model in a few hundred steps. There is no cheap undo.

The idea behind parameter-efficient fine-tuning

Here is the observation that started it. Freeze the pretrained weights entirely. Add a small number of new parameters. Train only those. If the new parameters are placed well, the resulting model performs about as well as full fine-tuning on the target task — while the number of trainable parameters drops by two or three orders of magnitude.

The underlying claim is that adapting a pretrained model to a downstream task does not require a full-rank, unconstrained change to every matrix. The pretrained model already has the representations; adaptation is mostly about steering them. Steering is a low-dimensional operation. You do not need 6.74 billion knobs to tell a model "answer in the style of a radiology report".

Pretraining teaches the model what language is. Fine-tuning usually just tells it which part of what it already knows to emphasise — and that instruction is far smaller than the knowledge itself.

The four families of PEFT method

LoRA — low-rank adaptation

Instead of learning a full update matrix for a weight W of shape 4096 × 4096, learn two thin matrices whose product has the same shape. With rank rr, you learn A∈Rr×4096A \in \mathbb{R}^{r \times 4096} and B∈R4096×rB \in \mathbb{R}^{4096 \times r}, and the effective weight becomes W+BAW + BA. The forward pass is h=Wx+BAxh = Wx + BAx.

Count it. A full update to that one matrix is 4096 × 4096 = 16,777,216 parameters. With r=8r = 8, LoRA learns 8 × 4096 + 4096 × 8 = 65,536 parameters. That is 0.39% of the full update, for one matrix.

Apply it to the query and value projections of all 32 layers of Llama-2-7B, at r=16r = 16:

Text
Per matrix:   16 x 4096 (A) + 4096 x 16 (B) = 131,072Per layer:    131,072 x 2 matrices (q, v)   = 262,144All layers:   262,144 x 32                  = 8,388,6088,388,608 / 6,738,415,616 = 0.124% of the model

Eight million trainable parameters instead of 6.74 billion. Because only those eight million need gradients and optimiser state, the 16-bytes-per-parameter rule now applies to 8.39M rather than 6.74B: about 134 MB instead of 108 GB. The frozen base weights still occupy 13.5 GB in BF16, but they are read-only.

Prefix tuning

Prepend a set of trainable vectors to the key and value sequences at every attention layer. The input text is untouched; the model simply sees some extra "virtual tokens" it can attend to, and those vectors are optimised. For Llama-2-7B with 32 layers, hidden size 4096 and 20 prefix vectors: 2 × 32 × 20 × 4096 = 5,242,880 parameters, about 0.078% of the model.

Prefix tuning is expressive because it acts at every layer, but it eats context window — those 20 virtual tokens occupy attention positions on every single forward pass, forever.

Prompt tuning

The minimal version: trainable vectors at the input embedding layer only, nowhere else. For 20 soft-prompt tokens at hidden size 4096, that is 20 × 4096 = 81,920 parameters — 0.0012% of the model. Astonishingly small. The catch is that prompt tuning only becomes competitive at very large model scales (roughly 10B+ parameters); on smaller models it consistently underperforms because a single layer of influence is not enough leverage.

Adapters

Insert small bottleneck modules between existing layers: project the hidden state down from 4096 to a bottleneck dimension, apply a non-linearity, project back up, and add residually. With bottleneck 64 and two adapters per layer (one after attention, one after the feed-forward block): 2 × (4096 × 64 + 64 × 4096) × 32 layers = 33,554,432 parameters, about 0.50%.

Adapters work well but have a structural disadvantage: they are sequential new layers. Every forward pass at inference must run through them, adding measurable latency. LoRA's adapters are parallel to an existing matrix and can be folded into it after training, so they add zero latency once merged.

Side by side

MethodTrainable params (7B model)% of modelInference latencyNotes
Full fine-tuning6,738,415,616100%None addedNeeds ~108 GB optimiser state
Adapters (bottleneck 64)33,554,4320.50%Added, permanentExtra sequential layers
LoRA (r=16, q+v)8,388,6080.124%Zero if mergedFoldable into base weights
Prefix tuning (20 vectors)5,242,8800.078%Added, permanentConsumes context length
Prompt tuning (20 tokens)81,9200.0012%Small, permanentOnly works at large scale

The storage argument, which is often the decisive one

A full fine-tune produces a 13.5 GB checkpoint. A LoRA adapter at r=16r = 16 produces 8,388,608 × 2 bytes = 16.8 MB in BF16. Go back to the thirty medical specialities:

Full fine-tuningLoRA (r=16)
Artefact per speciality13.5 GB16.8 MB
30 specialities on disk405 GB504 MB
GPU memory to serve all 3030 × 13.5 = 405 GB13.5 GB base + 504 MB
Cost to add speciality 31A new deploymentLoad a 17 MB file

The strongest argument for LoRA is often not training memory but serving economics: one base model in memory, hundreds of swappable 17 MB adapters on top of it.

What PEFT gives up

PEFT is not free. Be honest about the trade-offs:

  • Ceiling on how much can change. If your target domain is genuinely far from pretraining — a new script, a new language, protein sequences — a rank-16 update in a handful of matrices is not enough capacity, and full fine-tuning (or continued pretraining) will beat it.
  • New hyperparameters. Rank, scaling factor, and which modules to target are three more choices to get wrong.
  • Quality gap on some tasks. Usually 0-2% on task metrics, but "usually" is doing work in that sentence. Measure it.
  • Serving complexity if you do not merge. Unmerged adapters mean an extra code path in production.

Choosing between them

SituationChooseWhy
Under ~50,000 training examples, style/format/domain adaptationLoRACapacity is not the bottleneck; efficiency is
Many task variants sharing one base modelLoRAAdapter swapping is the entire point
One GPU, 24 GB or lessLoRA, quantised baseFull fine-tuning simply will not fit
Millions of examples, genuinely new domain or languageFull fine-tuningYou need full-rank capacity
Building a new base model to serve as a foundationFull fine-tuningThe output is meant to be the new starting point
Last 0.5% of accuracy matters more than any costFull fine-tuningIt is still the ceiling
Extremely large model (65B+) on limited hardwareLoRA on a 4-bit baseThe only option that fits at all

Three scenarios worked through

Medical domain adaptation. 4,000 radiology summaries, one 40 GB GPU, and a hard requirement that the model stay good at general instruction-following. LoRA at r=16r = 16. Full fine-tuning does not fit on the hardware, and with only 4,000 examples it would badly overfit and forget. The frozen base is a built-in regulariser: 99.88% of the model literally cannot drift.

Customer support chatbot, 40 enterprise tenants. Each tenant needs its own tone, product vocabulary and escalation policy. LoRA, decisively. 40 adapters at 16.8 MB each is 672 MB total, served from one 13.5 GB base. Full fine-tuning would mean 40 separate 13.5 GB models and 40 deployments — around 540 GB of GPU memory versus 14 GB.

Research lab training a new base model. 200 billion tokens of a language poorly represented in the original pretraining mix, 64 A100s available, and the output is meant to be a foundation others build on. Full fine-tuning (really continued pretraining). The dataset is large enough to avoid overfitting, the hardware exists, and a low-rank constraint would genuinely cap how much the model can learn.

What this means when you sit down to build

Start by computing two numbers before you write any training code. First, parameter count × 16 bytes — that is your full fine-tuning memory floor, and if it exceeds your GPU, the decision is already made for you. Second, number of distinct models you will eventually need to serve × checkpoint size — that is your storage and serving bill, and it is the number that surprises people six months in.

The failure mode to avoid is reaching for full fine-tuning by default because it feels like the "real" version. With a few thousand examples it is usually the worse choice on quality as well as cost: the model has enough freedom to memorise your small dataset and enough freedom to forget everything else. Teams routinely ship a LoRA fine-tune that beats their own full fine-tune, and are surprised, and should not be.

The opposite failure is treating PEFT as universal. If evaluation shows your LoRA run plateauing well short of the target and increasing rank keeps helping right up to r=256r = 256, that is the model telling you the low-rank assumption does not hold for your task. Believe it, and stop fighting the constraint.