Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is LoRA and QLoRA?
What you need to know
Why full fine-tuning is expensive
Full fine-tuning updates every weight. With the AdamW optimiser and mixed precision, you need roughly 16 bytes per parameter for weights, gradients and optimiser state. For an 8B model that is about 128 GB before activations — several large GPUs — and every task produces a new 16 GB copy of the model.
LoRA with numbers
One attention projection is 4,096 × 4,096 = 16.8 million weights. With r = 8:
A: 8 x 4,096 = 32,768B: 4,096 x 8 = 32,768Trainable = 65,536 (0.39% of the matrix)Apply LoRA to the query and value projections in 32 layers: 64 matrices × 65,536 ≈ 4.2 million trainable parameters, about 0.05% of an 8B model. Stored in 16-bit, the adapter is about 8 MB.
The layer computes W·x + (alpha / r)·B·A·x. B starts at zero, so training begins exactly at the base model. After training you can merge: W_new = W + (alpha / r)·B·A.
QLoRA additions
- 4-bit NF4 — the frozen base weights are stored in a 4-bit format designed for normally distributed weights; an 8B model drops from about 16 GB to about 5 GB.
- Double quantization — the quantization constants are themselves quantized to save more memory.
- Paged optimisers — optimiser state can spill to CPU memory to survive memory spikes.
- Compute in 16-bit — weights are de-quantized on the fly for each matrix multiply; the LoRA adapters train in 16-bit.
The QLoRA paper reported fine-tuning a 65B model on a single 48 GB GPU.
1import torch2from transformers import AutoModelForCausalLM, BitsAndBytesConfig3from peft import LoraConfig, get_peft_model45bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",6 bnb_4bit_use_double_quant=True,7 bnb_4bit_compute_dtype=torch.bfloat16)8model = AutoModelForCausalLM.from_pretrained("your-8b-model", quantization_config=bnb)910lora = LoraConfig(r=8, lora_alpha=16, lora_dropout=0.05,11 target_modules=["q_proj", "v_proj"], task_type="CAUSAL_LM")12model = get_peft_model(model, lora)13model.print_trainable_parameters()The first block loads the base model in 4-bit (QLoRA); the second attaches rank-8 adapters to the query and value projections. Remove the quantization_config and it becomes plain LoRA.
Trade-offs
- LoRA matches full fine-tuning well for style, format, tone and narrow domain behaviour.
- It is weaker at adding large amounts of new knowledge or new skills such as a new programming language; retrieval is usually better for facts.
- QLoRA trains somewhat slower than LoRA because of de-quantization, and the final model is usually served in 16-bit or re-quantized.
A real-life example
A law firm wants its summariser to always output the firm's format: parties, key dates, obligations, termination terms, and a risk flag, in formal Indian legal English. Prompting gets the format right about 85% of the time on long contracts.
They QLoRA-train an open 8B model on 3,000 contract–summary pairs written by their associates: one 24 GB GPU, r = 16, about 4 hours. The format is now right on almost every test contract, and the adapter is 20 MB. Later they train separate adapters for two practice areas and serve all of them on one base model, loading the right adapter per request.
Follow-up questions to expect
- "How do you choose the rank r?" — Start with 8 or 16; raise it if the task is complex and validation keeps improving. Higher rank means more capacity and memory.
- "Which layers should get LoRA?" — Attention projections are the classic choice; adding the MLP projections usually helps quality for a modest cost.
- "Does LoRA add inference latency?" — Not after merging. Unmerged adapters add a small cost but let you swap tasks per request.