Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is QLoRA, and how does 4-bit training make consumer-hardware fine-tuning possible?
What you need to know
Bytes per parameter
| Format | Bytes per parameter | 8B model weights |
|---|---|---|
| fp32 | 4 | 32 GB |
| bf16 / fp16 | 2 | 16 GB |
| 8-bit | 1 | 8 GB |
| 4-bit (NF4) | 0.5 | 4 GB, plus scale factors |
The three QLoRA ideas
- 4-bit NormalFloat (NF4). Pretrained weights are roughly bell-shaped (normally distributed), so NF4 puts its 16 levels where weights are most common instead of spacing them evenly. Section 6 covers NF4 in depth.
- Double quantisation. Weights are quantised in blocks of 64, each with its own fp32 scale: 32 bits ÷ 64 = 0.5 extra bits per parameter. QLoRA quantises those scales too (to 8-bit, in blocks of 256), cutting the overhead to about 0.127 bits — saving roughly 0.37 bits per parameter, about 370 MB on an 8B model.
- Paged optimizers. Optimizer state can spill to CPU memory during a memory spike (for example, one very long sequence) instead of crashing with out-of-memory. In Hugging Face this is
optim="paged_adamw_8bit"or"paged_adamw_32bit".
Why an 8B model is about 5–6 GB, not 4 GB
bitsandbytes quantises the linear layers but usually keeps the token embeddings and the output layer in bf16. On Llama-3.1-8B those two are about 1.05B parameters, or 2.1 GB. The other 7B parameters at about 0.52 bytes each add about 3.6 GB. Total: roughly 5.7 GB.
Code
1import torch2from transformers import AutoModelForCausalLM, BitsAndBytesConfig3from peft import prepare_model_for_kbit_training45bnb = BitsAndBytesConfig(6 load_in_4bit=True, bnb_4bit_quant_type="nf4",7 bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)8base = AutoModelForCausalLM.from_pretrained(9 "meta-llama/Llama-3.1-8B-Instruct", quantization_config=bnb)10base = prepare_model_for_kbit_training(base) # enables checkpointing, casts normsAfter this you add a LoraConfig exactly as in plain LoRA. bnb_4bit_compute_dtype is the dtype the weights are dequantised to for each matrix multiply. bitsandbytes 4-bit training is mainly used on NVIDIA GPUs. Section 4 covers the config options in more detail.
The trade-offs
- Speed. Dequantising every layer on every step makes QLoRA noticeably slower per step than bf16 LoRA.
- Quality. The QLoRA paper reported matching 16-bit fine-tuning on its benchmarks. In practice, measure it on your own task.
- Merging. Merge the adapter into a bf16 copy of the base, not into the 4-bit weights, then re-quantise for serving if needed.
A real-life example
A hospital group wants a summariser for radiology reports. Patient data may not leave the hospital, and the team has one workstation with a 24 GB RTX 4090.
bf16 LoRA: base 16 GB + adapter ~0.7 GB + activations -> too tight on 24 GBQLoRA: base ~5.7 GB + adapter ~0.7 GB + activations -> fits, with room for 4,096-token reports using gradient checkpointingThey train with QLoRA on 2,000 radiologist-edited summaries. Each epoch takes longer than it would with bf16 LoRA on a bigger GPU, but it runs on the hardware they are allowed to use.
Follow-up questions to expect
- "Are the LoRA adapters 4-bit too?" — No. Adapters are trained in bf16 or fp32. Only the frozen base is 4-bit.
- "Why NF4 and not plain int4?" — Evenly spaced int4 levels waste many levels on rare large values. NF4 matches the bell-shaped weight distribution, so the error is lower for the same 4 bits.
- "When would you use plain LoRA instead?" — When memory is not the constraint. bf16 LoRA is faster per step and avoids quantisation error.