Course Content
LLMs Deep Dive
10 sections · 40 lessons
How does Parameter-Efficient Fine-Tuning (PEFT) prevent catastrophic forgetting?
What you need to know
What each PEFT method trains
| Method | What is trained | Typical size |
|---|---|---|
| LoRA | Low-rank matrices B·A added to chosen weight matrices | 0.05–1% of parameters |
| Adapters | Small bottleneck layers inserted after sub-layers | 1–5% |
| Prefix / prompt tuning | Learned "virtual token" vectors prepended to the input or to each layer | Under 0.1% |
| IA3 | Learned scaling vectors on activations | Under 0.1% |
In all of them, the original weights W are never written to.
Four reasons forgetting is reduced
Full fine-tuning
- Every weight can move
- Old skills overwritten in place
- Undo requires the saved original model
- One task per model copy
LoRA / PEFT
- Base weights frozen; only the adapter moves
- Low rank limits how much behaviour can change
- Remove the adapter and the original is back exactly
- One base, many small adapters, isolated per task
Research published in 2024 comparing the two reached a memorable summary: LoRA "learns less and forgets less" than full fine-tuning — smaller gains on hard new domains such as maths or code, but better retention of the base model's other skills.
Where the protection ends
The frozen base protects the model when the adapter is off. When it is on, the output is W·x + B·A·x, and the adapter's term can still push general behaviour in the wrong direction. Causes:
- High rank (e.g. 256) applied to all layers — close to full fine-tuning.
- High learning rate or many epochs on narrow data.
- Training data that directly contradicts general behaviour (for example, every answer in one language).
So you still mix data, train gently and run the regression suite.
A real-life example
A law firm serves three LoRA adapters on one base model: contract summaries, litigation brief drafting and a Hindi-to-English translation helper for court documents. When the litigation team retrains its adapter with a bad batch of data, only litigation outputs change; contract summaries and translations are untouched, because each request loads its own adapter and the base weights never changed.
But the litigation adapter itself, trained for 6 epochs at r = 128, started ignoring the "cite the paragraph number" instruction that the base model followed well. PEFT had not prevented forgetting within that adapter. Retraining at r = 16 for 2 epochs with 15% general instruction examples restored it.
Follow-up questions to expect
- "Can you combine several LoRA adapters?" — Yes, by adding or merging them, but they can interfere; it works best when the tasks are unrelated and is always worth evaluating.
- "Is prefix tuning as good as LoRA?" — Usually weaker on large tasks and more sensitive to settings; LoRA is the common default.
- "Why not always use PEFT?" — When you need the biggest possible gain on a large new domain, and have the compute, full fine-tuning can still win.