Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is QLoRA, and how does 4-bit training make consumer-hardware fine-tuning possible?


An 8B QLoRA run on one 24 GB cardLinear layersin NF4: ~3.6 GBEmbeddings + headin bf16: ~2.1 GBLoRA adapter +Adam: ~0.7 GBCheckpointedactivationsHeadroom; pagedoptimizer spilltopbottomThe same base in bf16 would take 16 GB before anything else.
Only the frozen base is 4-bit — the adapters and every matrix multiply still run in bf16, which is why quality holds up.

What you need to know

Bytes per parameter

FormatBytes per parameter8B model weights
fp32432 GB
bf16 / fp16216 GB
8-bit18 GB
4-bit (NF4)0.54 GB, plus scale factors

The three QLoRA ideas

  • 4-bit NormalFloat (NF4). Pretrained weights are roughly bell-shaped (normally distributed), so NF4 puts its 16 levels where weights are most common instead of spacing them evenly. Section 6 covers NF4 in depth.
  • Double quantisation. Weights are quantised in blocks of 64, each with its own fp32 scale: 32 bits ÷ 64 = 0.5 extra bits per parameter. QLoRA quantises those scales too (to 8-bit, in blocks of 256), cutting the overhead to about 0.127 bits — saving roughly 0.37 bits per parameter, about 370 MB on an 8B model.
  • Paged optimizers. Optimizer state can spill to CPU memory during a memory spike (for example, one very long sequence) instead of crashing with out-of-memory. In Hugging Face this is optim="paged_adamw_8bit" or "paged_adamw_32bit".

Why an 8B model is about 5–6 GB, not 4 GB

bitsandbytes quantises the linear layers but usually keeps the token embeddings and the output layer in bf16. On Llama-3.1-8B those two are about 1.05B parameters, or 2.1 GB. The other 7B parameters at about 0.52 bytes each add about 3.6 GB. Total: roughly 5.7 GB.

Code

Python
import torchfrom transformers import AutoModelForCausalLM, BitsAndBytesConfigfrom peft import prepare_model_for_kbit_trainingbnb = BitsAndBytesConfig(    load_in_4bit=True, bnb_4bit_quant_type="nf4",    bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)base = AutoModelForCausalLM.from_pretrained(    "meta-llama/Llama-3.1-8B-Instruct", quantization_config=bnb)base = prepare_model_for_kbit_training(base)   # enables checkpointing, casts norms

After this you add a LoraConfig exactly as in plain LoRA. bnb_4bit_compute_dtype is the dtype the weights are dequantised to for each matrix multiply. bitsandbytes 4-bit training is mainly used on NVIDIA GPUs. Section 4 covers the config options in more detail.

The trade-offs

  • Speed. Dequantising every layer on every step makes QLoRA noticeably slower per step than bf16 LoRA.
  • Quality. The QLoRA paper reported matching 16-bit fine-tuning on its benchmarks. In practice, measure it on your own task.
  • Merging. Merge the adapter into a bf16 copy of the base, not into the 4-bit weights, then re-quantise for serving if needed.

A real-life example

A hospital group wants a summariser for radiology reports. Patient data may not leave the hospital, and the team has one workstation with a 24 GB RTX 4090.

Text
bf16 LoRA:  base 16 GB + adapter ~0.7 GB + activations  -> too tight on 24 GBQLoRA:      base ~5.7 GB + adapter ~0.7 GB + activations -> fits, with room            for 4,096-token reports using gradient checkpointing

They train with QLoRA on 2,000 radiologist-edited summaries. Each epoch takes longer than it would with bf16 LoRA on a bigger GPU, but it runs on the hardware they are allowed to use.

Follow-up questions to expect

  • "Are the LoRA adapters 4-bit too?" — No. Adapters are trained in bf16 or fp32. Only the frozen base is 4-bit.
  • "Why NF4 and not plain int4?" — Evenly spaced int4 levels waste many levels on rare large values. NF4 matches the bell-shaped weight distribution, so the error is lower for the same 4 bits.
  • "When would you use plain LoRA instead?" — When memory is not the constraint. bf16 LoRA is faster per step and avoids quantisation error.