Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Why is LoRA effective? What intuition or theory explains it?
What you need to know
What "low rank" means in plain words
A rank-1 update is one rule: "when the input points in this direction, push the output in that direction." A rank-16 update is 16 such rules per matrix. Fine-tuning for a format, a tone or a classification boundary needs a few such rules, not millions of independent weight changes.
The evidence
- Intrinsic dimension (2020). Aghajanyan and colleagues showed that fine-tuning large pretrained models works even when the update is restricted to a small random subspace, and that larger models need fewer dimensions. Pretraining gives a good starting point from which small moves go a long way.
- The LoRA paper (Hu et al., 2021). Low-rank updates, even at very small ranks, matched full fine-tuning on many tasks with a tiny fraction of the trainable parameters.
- "LoRA Learns Less and Forgets Less" (2024). On large code and maths training, LoRA fell short of full fine-tuning, but it kept more of the base model's other skills. The low rank acts as a limit on how far the model can move.
- Later studies (2025). Reports such as Thinking Machines' "LoRA Without Regret" found LoRA matches full fine-tuning on small-to-medium post-training datasets when it is applied to all layers, including the MLP, with a properly tuned learning rate, and that RL fine-tuning needs very little capacity.
Why RL needs so little rank
In supervised fine-tuning, every target token carries information. In RL, each whole answer earns one reward number, so the training signal per example is small. A small adapter has enough room to store it.
Rank is a capacity dial
| Task | Typical rank |
|---|---|
| Tone, format, simple classification | 8–16 |
| Multi-task, harder reasoning, long outputs | 32–64 |
| Large domain shift or new language | LoRA struggles; consider continued pretraining or full fine-tuning |
If raising the rank from 16 to 64 does not improve validation results, capacity is not your bottleneck. Look at your data.
A real-life example
A team building the Hindi support assistant tries LoRA at ranks 8, 16 and 64 on 3,000 chats. On their 300-chat test set, all three score within a point of each other on tone and correctness, so they ship rank 16 — the smallest adapter that is reliably good.
Later they try to add Odia, which the base model saw little of in pretraining, using the same LoRA setup on a few thousand chats. Replies are fluent in Hindi but broken in Odia. That is the limit the theory predicts: a new language is not a small re-weighting. They switch to continued pretraining on Odia text first, then LoRA for the support behaviour.
Follow-up questions to expect
- "Is the low-rank idea proven?" — It is an empirical finding, not a theorem. It holds well for adapting behaviour and weakly for learning lots of new knowledge.
- "Why does LoRA forget less?" — The update is small and constrained, and the original weights are untouched, so there is less room to overwrite old skills. It is less forgetting, not zero.
- "Why not always use a very high rank?" — More memory, more risk of overfitting a small dataset, and high ranks need rescaling (rsLoRA) to train stably.