Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Contrast full fine-tuning with parameter-efficient fine-tuning (PEFT)


Training memory for an 8B model, before activations16 GB16 GB96 GBabout 128 GB16 GB frozen0.17 GB0.34 GBabout 17 GBWeightsGradientsMaster + AdamTotalFull fine-tuningLoRA r=16, all-linearLoRA trains 41.9M of 8.03B parameters.
LoRA removes the optimizer and gradient bill, but the frozen 16 GB base is still there — it cuts memory about 7 times, not 200 times.

What you need to know

Where the 16 bytes per parameter come from

In full fine-tuning, every parameter needs its weight, its gradient, and the optimizer's memory. Adam keeps two running averages per parameter (momentum and variance), and mixed-precision training keeps an fp32 "master" copy of the weights.

ItemBytes per parameterLlama-3.1-8B (8.03B params)
Weights (bf16)216 GB
Gradients (bf16)216 GB
fp32 master weights432 GB
Adam momentum + variance (fp32)864 GB
Total16about 128 GB, plus activations

That does not fit on one 80 GB GPU. You need several GPUs and a sharding method such as FSDP or DeepSpeed ZeRO.

The same model with LoRA

LoRA on all linear layers of Llama-3.1-8B at rank 16 adds 41,943,040 trainable parameters (0.52%). Only those get gradients and optimizer state:

Text
frozen base (bf16):        8.03B x 2 bytes   = about 16 GBadapter (fp32, with Adam): 41.9M x 16 bytes  = about 0.7 GBtotal before activations:                    about 17 GB

With gradient checkpointing and moderate sequence lengths, that fits one 24–48 GB GPU. With QLoRA (a 4-bit base) the frozen part shrinks to about 5–6 GB.

The PEFT family at a glance

MethodWhat it trainsNote
LoRALow-rank update matrices beside frozen weightsThe default; mergeable, no inference cost after merging
QLoRALoRA on a 4-bit frozen baseLeast memory; slower steps
DoRALoRA plus a learned magnitude per outputOften a little better at low rank; slower
Bottleneck adaptersSmall MLP blocks inserted in each layerThe original adapter method; adds latency
Prompt / prefix tuningLearned vectors fed as input or as keys/valuesTiny, but usually weaker (section 2)

When full fine-tuning is still right

  • A real distribution shift: a new language or script, a new modality, very unusual notation.
  • Continued pretraining on hundreds of millions of tokens or more.
  • You serve exactly one model, you measured LoRA properly, and it plateaus below target.

A real-life example

A diagnostics chain wants three summarisers for its doctors: radiology reports, pathology reports and discharge notes. Its budget is one 48 GB GPU.

  • Full fine-tuning an 8B model needs about 128 GB before activations, so it would rent a multi-GPU node for each run, and store three separate 16 GB checkpoints.
  • LoRA fits on the one GPU. Each of the three adapters is about 84 MB in bf16 (41.9M × 2 bytes), all three run over one shared base model, and rolling back a bad adapter means switching a file.

After training, the base model's general skills are unchanged when the adapter is switched off — useful when a doctor also asks the assistant general questions.

Follow-up questions to expect

  • "Does PEFT reduce activation memory?" — Not much. Activations still flow through the whole network in the forward and backward pass. Use gradient checkpointing, shorter sequences or smaller micro-batches for that.
  • "Is PEFT faster?" — Somewhat. You skip weight gradients and optimizer updates for frozen weights, but the forward and backward passes still run through every layer, so do not expect a 100× speed-up.
  • "Is LoRA always worse than full fine-tuning?" — No. On small and medium post-training datasets, LoRA on all linear layers with a tuned learning rate often matches it. The gap shows on large datasets and big distribution shifts.