Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is LoRA (Low-Rank Adaptation) and what changes during training?
What you need to know
The formula in words
original: h = W xwith LoRA: h = W x + (alpha / r) * B (A x)W: d_out x d_in (frozen)A: r x d_in (trained, starts small and random)B: d_out x r (trained, starts at zero)A squeezes the input down to r numbers, and B expands them back to the output size. Because B starts at zero, B A x is zero at step 0, so the model starts exactly as the base model. A is random, so the gradient reaching B is not zero and learning can start. (If both started at zero, neither would ever receive a gradient.)
Worked arithmetic on a 4096 × 4096 matrix
LoRA parameters = r × (d_in + d_out) = r × 8,192. The full matrix has 16,777,216.
| Rank r | LoRA params | Share of the matrix |
|---|---|---|
| 4 | 32,768 | 0.20% |
| 8 | 65,536 | 0.39% |
| 16 | 131,072 | 0.78% |
| 64 | 524,288 | 3.1% |
Across a whole model the numbers add up. On Llama-3.1-8B at r = 16, targeting only the four attention projections gives 13.6M trainable parameters (0.17%); targeting all seven linear layers per block, including the MLP, gives 41.9M (0.52%).
Code (peft 0.21, transformers 5.x)
1import torch2from transformers import AutoModelForCausalLM3from peft import LoraConfig, get_peft_model45base = AutoModelForCausalLM.from_pretrained(6 "meta-llama/Llama-3.1-8B-Instruct", dtype=torch.bfloat16)7config = LoraConfig(8 r=16, lora_alpha=32, lora_dropout=0.05,9 target_modules="all-linear", # q,k,v,o and gate,up,down10 task_type="CAUSAL_LM")11model = get_peft_model(base, config)12model.print_trainable_parameters()13# trainable params: 41,943,040 || all params: 8,072,204,288 || trainable%: 0.5196r sets capacity. lora_alpha sets the strength: the update is multiplied by alpha / r, here 2. Common defaults are alpha = r or alpha = 2r. use_rslora=True scales by alpha / √r instead, which keeps training stable when you use large ranks. use_dora=True switches to DoRA, which also learns a magnitude per output channel (1.4M extra parameters on this model).
After training
model.merge_and_unload() adds (alpha / r) * B A into W and removes the adapter, so inference is as fast as the base model. Or keep adapters separate and serve many of them over one base (section 4 covers serving).
A real-life example
A Mumbai law firm reviews hundreds of commercial contracts a month. It wants each clause labelled — indemnity, limitation of liability, arbitration seat, governing law, termination, confidentiality — with a one-line risk note.
The team trains a rank-16, all-linear LoRA adapter on 6,000 clauses labelled by associates. The adapter file is about 84 MB, while the base model is 16 GB. Later, the firm wants a second behaviour — drafting fallback clauses — so it trains a second adapter over the same base instead of a second model. The frozen base means the model's general English and reasoning stay intact, and the firm can prove which adapter version produced any label.
Follow-up questions to expect
- "Which modules should you target?" — All linear layers usually beat attention-only at the same budget, because much of the knowledge lives in the MLP layers.
target_modules="all-linear"does this in PEFT. - "What does lora_alpha do?" — It scales the update. If you double
rbut keep alpha fixed, each direction's contribution halves, so re-tune alpha or the learning rate when you change rank. - "Does LoRA add inference latency?" — Only if unmerged: an extra small matrix multiply per targeted layer. Merged, it adds nothing.
- "What learning rate?" — LoRA typically uses a learning rate around 10 times higher than full fine-tuning, for example 1e-4 to 2e-4 instead of 1e-5 to 2e-5.