Course Content
Fine-Tuning LLMs with LoRA, QLoRA and PEFT
4 sections · 10 lessons
LoRA Architecture — Low-Rank Adapters
Open a Llama-2-7B checkpoint and look at one weight matrix — say the query projection in layer 12. It is 4096 rows by 4096 columns: 16,777,216 numbers. There are 32 layers, each with four such attention matrices plus three larger feed-forward matrices. Altogether, 6,738,415,616 parameters.
Now suppose you want the model to write in the style of your company's release notes. You have 3,000 examples. Full fine-tuning says: adjust all 6.74 billion numbers. Take that seriously for a moment — you are proposing to estimate 6.74 billion free parameters from 3,000 examples, roughly 2.2 million parameters per training example.
Stated that way it sounds absurd, and the memory bill agrees: AdamW in mixed precision costs about 16 bytes per parameter, so 6.74B × 16 = 108 GB of weights, gradients and optimiser state before a single activation is stored. That is more than an 80 GB card holds.
LoRA starts from a suspicion: the change you need is not 6.74 billion numbers wide. It is far narrower, and if you can express that narrowness in the architecture, everything else — memory, storage, deployment — collapses with it.
The low-rank hypothesis
Every matrix has a rank: the number of genuinely independent directions it acts along. A 4096 × 4096 matrix can have rank up to 4096, but it need not. If a matrix has rank 8, then despite having 16.8 million entries, all of its behaviour is captured by 8 independent directions — and it can be written exactly as the product of a 4096 × 8 matrix and an 8 × 4096 matrix.
LoRA's claim is not about the pretrained weights. It is about the update. When you fine-tune a pretrained model on a downstream task, the difference between the final and initial weights, ΔW=Wfinal−W0, turns out to be approximately low rank. Take the singular values of that difference matrix and most of them are near zero; a handful carry nearly all the energy.
That makes intuitive sense. Pretraining spent trillions of tokens building general representations. Fine-tuning on 3,000 release notes is not rebuilding those representations — it is applying a comparatively simple, systematic adjustment. Simple systematic adjustments are low rank.
LoRA does not claim that language models are low rank. It claims that the change from a pretrained model to a fine-tuned one is — and that you can therefore parameterise the change directly, instead of parameterising the whole model again.
The mechanism, one layer at a time
The layer you already have
A linear layer computes:
with W0∈Rd×k. For a Llama-2-7B attention projection, d=k=4096. Full fine-tuning replaces W0 with W0+ΔW where every entry of ΔW is a free parameter — 16,777,216 of them for this one matrix.
The layer with an adapter
LoRA freezes W0 completely — no gradients, no optimiser state, never updated — and forces the update to factor through a bottleneck of width r:
The forward pass becomes:
Read it as two parallel paths. The frozen path W0x produces what the pretrained model would have produced. The adapter path squeezes x down from 4096 dimensions to r, then expands back to 4096, and the result is added on. Because BA is a product through an r-wide bottleneck, its rank can never exceed r — the constraint is structural, not a penalty term.
x (4096) | +----------+-----------+ | | W0 (frozen) A (r x 4096) -> trainable 4096 x 4096 | | z (r) | | | B (4096 x r) -> trainable | | | scale by alpha/r | | +----------+-----------+ | h (4096)Why A is random and B is zero
At initialisation, A is filled with small random values (Kaiming-style) and B is filled with exact zeros. That means BA=0, so at step 0 the adapted model computes precisely what the pretrained model computes. Training starts from the pretrained behaviour, not from a perturbed version of it.
Why not initialise both randomly? Because then BA=0 at step 0 and you inject structured noise into every adapted layer simultaneously. The first loss value spikes and the first few hundred steps are spent undoing damage you created.
Why not initialise both to zero? Work out the gradients. With h=W0x+BAx:
If A=0, then Ax=0 and ∂L/∂B=0. If B=0, then B⊤=0 and ∂L/∂A=0. Set both to zero and both gradients vanish forever — the adapter is dead and never learns anything. Exactly one of them must be non-zero, and the asymmetric choice (random A, zero B) gives you a zero-valued product with live gradients. On the first backward pass B moves off zero, and from step 2 onwards A has a non-zero gradient too.
Random A with zero B is the only initialisation that starts training from exactly the pretrained model while keeping both matrices learnable. Both-random breaks the starting point; both-zero breaks the gradients.
The scaling factor alpha
The α/r factor exists to decouple the strength of the adapter from its width. Without it, doubling the rank roughly doubles the magnitude of BAx, and a learning rate tuned at r=8 would be badly wrong at r=16. With the factor, the effective contribution is normalised and you can sweep rank without re-tuning everything.
| alpha | r | Scale (alpha/r) | Effect |
|---|---|---|---|
| 8 | 8 | 1.0 | Conservative; adapter contributes gently |
| 16 | 8 | 2.0 | Common default; adapter has real influence |
| 16 | 16 | 1.0 | More capacity, same strength |
| 32 | 16 | 2.0 | The most widely used pairing |
| 64 | 16 | 4.0 | Aggressive; risks instability and forgetting |
The useful mental model: r is how many independent directions the adapter can express, and α/r is how loudly it speaks. A common convention is α=2r, which fixes the scale at 2.0 across every rank you try. Alpha is not a learning rate — it multiplies the adapter output on every forward pass, at training and at inference alike, so changing it after training changes the model's behaviour.
Where to put adapters inside a transformer
Each transformer block in Llama-2-7B contains seven weight matrices worth adapting:
| Module | Shape (7B) | Parameters | Role |
|---|---|---|---|
q_proj | 4096 × 4096 | 16,777,216 | What each token looks for |
k_proj | 4096 × 4096 | 16,777,216 | What each token offers |
v_proj | 4096 × 4096 | 16,777,216 | What gets carried forward |
o_proj | 4096 × 4096 | 16,777,216 | How heads are recombined |
gate_proj | 11008 × 4096 | 45,088,768 | Feed-forward gating |
up_proj | 11008 × 4096 | 45,088,768 | Feed-forward expansion |
down_proj | 4096 × 11008 | 45,088,768 | Feed-forward contraction |
The original LoRA work adapted only q_proj and v_proj, and reported that this was enough for the tasks studied. Practice has since shifted: for anything beyond mild style adaptation, targeting all seven linear modules generally performs better at the same total parameter budget than concentrating the budget on two of them. A useful rule is to prefer more modules at lower rank over fewer modules at high rank.
Layer norms are not adapted — they hold a few thousand parameters and are usually left frozen or made fully trainable, not low-rank. Embedding and output-head matrices are usually left alone unless you are adding new tokens, in which case they must be trained directly.
The arithmetic that makes the case
For a matrix of shape d×k, a LoRA adapter of rank r adds r(d+k) parameters. Everything follows from that one expression.
One matrix
q_proj: 4096 x 4096 = 16,777,216 parameters in the full updateLoRA r=8: 8 x (4096 + 4096) = 65,536 -> 0.391% of the full updateLoRA r=16: 16 x (4096 + 4096) = 131,072 -> 0.781%LoRA r=64: 64 x (4096 + 4096) = 524,288 -> 3.125%The whole model, q and v only
Two 4096 × 4096 matrices per layer, 32 layers, so the total is r×8192×2×32=r×524,288:
| Rank r | Trainable parameters | % of 6.74B | Adapter file (BF16) | Optimiser state (16 B/param) |
|---|---|---|---|---|
| 4 | 2,097,152 | 0.031% | 4.2 MB | 34 MB |
| 8 | 4,194,304 | 0.062% | 8.4 MB | 67 MB |
| 16 | 8,388,608 | 0.124% | 16.8 MB | 134 MB |
| 32 | 16,777,216 | 0.249% | 33.6 MB | 268 MB |
| 64 | 33,554,432 | 0.498% | 67.1 MB | 537 MB |
| Full FT | 6,738,415,616 | 100% | 13.5 GB | 108 GB |
The whole model, all seven modules
Per layer, at rank r:
Attention (4 matrices, each 4096 x 4096): 4 x r x (4096 + 4096) = r x 32,768Feed-forward (3 matrices, each pairing 4096 with 11008): 3 x r x (11008 + 4096) = r x 45,312Per layer total: r x 78,080All 32 layers: r x 2,498,560At r = 16: 39,976,960 trainable parameters = 0.593% of the modelForty million trainable parameters. Compare the memory ledger:
| Component | Full fine-tuning | LoRA r=16, all modules |
|---|---|---|
| Base weights (BF16) | 13.5 GB (trainable) | 13.5 GB (frozen) |
| Gradients | 13.5 GB | 0.16 GB |
| AdamW moments (2 × FP32) | 54.0 GB | 0.32 GB |
| FP32 master weights | 27.0 GB | 0.16 GB |
| Total (excluding activations) | 108 GB | 14.1 GB |
The base weights barely move — they are the same 13.5 GB either way. What disappears is the 94.5 GB of gradient and optimiser state, replaced by 640 MB. That is the entire trick: you do not pay 16 bytes per parameter for parameters you never update.
In code
1from transformers import AutoModelForCausalLM2from peft import LoraConfig, get_peft_model34model = AutoModelForCausalLM.from_pretrained(5 "meta-llama/Llama-2-7b-hf", dtype="bfloat16", device_map="auto"6)78config = LoraConfig(9 r=16,10 lora_alpha=32, # scale = 32/16 = 2.011 lora_dropout=0.05,12 bias="none",13 task_type="CAUSAL_LM",14 target_modules=["q_proj", "k_proj", "v_proj", "o_proj",15 "gate_proj", "up_proj", "down_proj"],16)1718model = get_peft_model(model, config)19model.print_trainable_parameters()20# trainable params: 39,976,960 || all params: 6,778,392,576 || trainable%: 0.5898Note that the reported "all params" is 6,778,392,576, not 6,738,415,616 — the difference is exactly the 39,976,960 adapter parameters that were just added. Always print this. If the trainable count is far from what your arithmetic predicts, a module name is wrong and some layers silently received no adapter at all.
Adapters are files you can move around
Because the base model is untouched, a trained adapter is a self-contained 17 MB artefact. That has consequences well beyond training memory.
Swapping at serve time
1from peft import PeftModel23base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf",4 dtype="bfloat16",5 device_map="auto")67legal = PeftModel.from_pretrained(base, "adapters/legal-summariser")8legal.set_adapter("default")910# Load a second adapter onto the same frozen base - no extra 13.5 GB11legal.load_adapter("adapters/medical-summariser", adapter_name="medical")12legal.set_adapter("medical")One 13.5 GB base in GPU memory, N adapters at 17 MB each. Fifty specialised models cost 13.5 GB + 850 MB, versus 675 GB if each were a full fine-tune.
Merging for deployment
When you only need one behaviour, fold the adapter into the base weights permanently:
legal.set_adapter("default") # merging uses the *active* adaptermerged = legal.merge_and_unload() # W0 <- W0 + (alpha/r) * B @ Amerged.save_pretrained("models/legal-7b-merged")The set_adapter line matters here because the previous block left medical active. merge_and_unload() folds in whichever adapter is active and throws the others away, so without it the "legal" checkpoint would quietly contain the medical adapter.
The result is an ordinary Llama-2-7B checkpoint with modified weights. It runs through any standard inference stack, and the adapter's runtime cost drops to exactly zero — the two matrices no longer exist, their effect having been added into W0. This is LoRA's advantage over bottleneck adapters, which insert genuinely new sequential layers that can never be folded away.
Where people get this wrong
"LoRA slows down inference." Unmerged, it adds two small matrix multiplications per adapted layer — typically 5-15% overhead. Merged, it adds nothing at all, because the maths is identical to a plain linear layer. If you are measuring a slowdown in production, you forgot to merge.
"Higher rank is always better." Rank sets an upper bound on the update's complexity; it does not force the model to use it. Past the point where the true update's rank is covered, extra rank adds parameters that mostly fit noise. On a 3,000-example dataset, r=64 frequently overfits where r=8 generalises. Sweep it; do not assume it.
"Alpha is a learning rate." It is a permanent output multiplier. Doubling alpha permanently doubles the adapter's influence at inference. Doubling the learning rate changes how fast the adapter learns and nothing about the final forward pass.
"LoRA can teach the model new facts." Only weakly. A rank-16 update spread across attention projections is well suited to changing how the model responds — format, tone, domain conventions, which of its existing abilities to apply. It is a poor vehicle for injecting a body of factual knowledge, which wants either retrieval or genuine continued pretraining.
"target_modules can be left at the default." Module names differ across architectures. Llama uses q_proj/k_proj/v_proj; GPT-2 uses a fused c_attn; Falcon uses query_key_value. Current PEFT raises an error if none of your names match, but a single wrong name in a list of seven is skipped without a word: training runs happily and those layers learn nothing. PEFT also accepts target_modules="all-linear", which targets every linear layer except the output head and saves you from spelling names at all. Either way, the trainable-parameter count is your check.
What this means when you configure a run
Before launching, compute the expected trainable parameter count by hand from r(d+k) summed over your target modules, then compare it to what print_trainable_parameters() reports. A mismatch means your configuration is not doing what you think, and it is far cheaper to catch that in ten seconds than after a four-hour run produces a model identical to the base.
Choose rank by dataset size rather than by ambition. A few thousand examples of style adaptation are well served by r=8 with all modules targeted. Tens of thousands of examples teaching genuinely new task behaviour justify r=32 or r=64. If a rank sweep shows the metric still improving at r=128, that is meaningful information: the update your task needs is not low rank, and you should be considering full fine-tuning rather than pushing rank further.
Finally, decide up front whether you will merge. If you are serving one behaviour, merge and forget the adapter exists. If you are serving many, keep them separate, keep the base loaded once, and treat adapters as configuration rather than as models — because at 17 MB apiece, that is what they are.