Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Contrast full fine-tuning with parameter-efficient fine-tuning (PEFT)
What you need to know
Where the 16 bytes per parameter come from
In full fine-tuning, every parameter needs its weight, its gradient, and the optimizer's memory. Adam keeps two running averages per parameter (momentum and variance), and mixed-precision training keeps an fp32 "master" copy of the weights.
| Item | Bytes per parameter | Llama-3.1-8B (8.03B params) |
|---|---|---|
| Weights (bf16) | 2 | 16 GB |
| Gradients (bf16) | 2 | 16 GB |
| fp32 master weights | 4 | 32 GB |
| Adam momentum + variance (fp32) | 8 | 64 GB |
| Total | 16 | about 128 GB, plus activations |
That does not fit on one 80 GB GPU. You need several GPUs and a sharding method such as FSDP or DeepSpeed ZeRO.
The same model with LoRA
LoRA on all linear layers of Llama-3.1-8B at rank 16 adds 41,943,040 trainable parameters (0.52%). Only those get gradients and optimizer state:
frozen base (bf16): 8.03B x 2 bytes = about 16 GBadapter (fp32, with Adam): 41.9M x 16 bytes = about 0.7 GBtotal before activations: about 17 GBWith gradient checkpointing and moderate sequence lengths, that fits one 24–48 GB GPU. With QLoRA (a 4-bit base) the frozen part shrinks to about 5–6 GB.
The PEFT family at a glance
| Method | What it trains | Note |
|---|---|---|
| LoRA | Low-rank update matrices beside frozen weights | The default; mergeable, no inference cost after merging |
| QLoRA | LoRA on a 4-bit frozen base | Least memory; slower steps |
| DoRA | LoRA plus a learned magnitude per output | Often a little better at low rank; slower |
| Bottleneck adapters | Small MLP blocks inserted in each layer | The original adapter method; adds latency |
| Prompt / prefix tuning | Learned vectors fed as input or as keys/values | Tiny, but usually weaker (section 2) |
When full fine-tuning is still right
- A real distribution shift: a new language or script, a new modality, very unusual notation.
- Continued pretraining on hundreds of millions of tokens or more.
- You serve exactly one model, you measured LoRA properly, and it plateaus below target.
A real-life example
A diagnostics chain wants three summarisers for its doctors: radiology reports, pathology reports and discharge notes. Its budget is one 48 GB GPU.
- Full fine-tuning an 8B model needs about 128 GB before activations, so it would rent a multi-GPU node for each run, and store three separate 16 GB checkpoints.
- LoRA fits on the one GPU. Each of the three adapters is about 84 MB in bf16 (41.9M × 2 bytes), all three run over one shared base model, and rolling back a bad adapter means switching a file.
After training, the base model's general skills are unchanged when the adapter is switched off — useful when a doctor also asks the assistant general questions.
Follow-up questions to expect
- "Does PEFT reduce activation memory?" — Not much. Activations still flow through the whole network in the forward and backward pass. Use gradient checkpointing, shorter sequences or smaller micro-batches for that.
- "Is PEFT faster?" — Somewhat. You skip weight gradients and optimizer updates for frozen weights, but the forward and backward passes still run through every layer, so do not expect a 100× speed-up.
- "Is LoRA always worse than full fine-tuning?" — No. On small and medium post-training datasets, LoRA on all linear layers with a tuned learning rate often matches it. The gap shows on large datasets and big distribution shifts.