Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Why is LoRA effective? What intuition or theory explains it?


What you need to know

What "low rank" means in plain words

A rank-1 update is one rule: "when the input points in this direction, push the output in that direction." A rank-16 update is 16 such rules per matrix. Fine-tuning for a format, a tone or a classification boundary needs a few such rules, not millions of independent weight changes.

The evidence

  • Intrinsic dimension (2020). Aghajanyan and colleagues showed that fine-tuning large pretrained models works even when the update is restricted to a small random subspace, and that larger models need fewer dimensions. Pretraining gives a good starting point from which small moves go a long way.
  • The LoRA paper (Hu et al., 2021). Low-rank updates, even at very small ranks, matched full fine-tuning on many tasks with a tiny fraction of the trainable parameters.
  • "LoRA Learns Less and Forgets Less" (2024). On large code and maths training, LoRA fell short of full fine-tuning, but it kept more of the base model's other skills. The low rank acts as a limit on how far the model can move.
  • Later studies (2025). Reports such as Thinking Machines' "LoRA Without Regret" found LoRA matches full fine-tuning on small-to-medium post-training datasets when it is applied to all layers, including the MLP, with a properly tuned learning rate, and that RL fine-tuning needs very little capacity.

Why RL needs so little rank

In supervised fine-tuning, every target token carries information. In RL, each whole answer earns one reward number, so the training signal per example is small. A small adapter has enough room to store it.

Rank is a capacity dial

TaskTypical rank
Tone, format, simple classification8–16
Multi-task, harder reasoning, long outputs32–64
Large domain shift or new languageLoRA struggles; consider continued pretraining or full fine-tuning

If raising the rank from 16 to 64 does not improve validation results, capacity is not your bottleneck. Look at your data.

A real-life example

A team building the Hindi support assistant tries LoRA at ranks 8, 16 and 64 on 3,000 chats. On their 300-chat test set, all three score within a point of each other on tone and correctness, so they ship rank 16 — the smallest adapter that is reliably good.

Later they try to add Odia, which the base model saw little of in pretraining, using the same LoRA setup on a few thousand chats. Replies are fluent in Hindi but broken in Odia. That is the limit the theory predicts: a new language is not a small re-weighting. They switch to continued pretraining on Odia text first, then LoRA for the support behaviour.

Follow-up questions to expect

  • "Is the low-rank idea proven?" — It is an empirical finding, not a theorem. It holds well for adapting behaviour and weakly for learning lots of new knowledge.
  • "Why does LoRA forget less?" — The update is small and constrained, and the original weights are untouched, so there is less room to overwrite old skills. It is less forgetting, not zero.
  • "Why not always use a very high rank?" — More memory, more risk of overfitting a small dataset, and high ranks need rescaling (rsLoRA) to train stably.