Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is catastrophic forgetting, and what techniques help avoid it while fine-tuning?
What you need to know
Why it happens
All skills share the same weights. During fine-tuning, the gradient only says "get better at this data". It has no information about maths, coding or politeness, so the update can move weights those skills depended on. The more the weights move — high learning rate, many epochs, full fine-tuning — the more is lost.
What gets forgotten
- General instruction following — ignoring formatting requests, answering in the wrong language.
- Reasoning and maths — especially when the new data has none.
- Safety behaviour. Research in 2023 (Qi et al.) showed that fine-tuning an aligned model, even on harmless data, can weaken its refusals. Re-test safety after every fine-tune.
- Other languages — a model tuned only on English support chats may get worse at Hindi.
Techniques that help
| Technique | How it helps | Cost |
|---|---|---|
| LoRA / PEFT | Base weights frozen; the update is small and can be switched off | Low |
| Low learning rate, 1–3 epochs | Less weight movement | Free |
| Replay (5–20% general data) | The loss keeps rewarding old skills | Some extra training time |
| KL or distillation term against the base | Penalises drifting from the base model's outputs on general prompts | Needs the base model in the loop |
| Weight averaging / lower adapter scale | Pulls the model part-way back toward the base | Small loss of task gain |
| Keep facts in retrieval | Less new knowledge pushed into weights | A retrieval system |
Replay data does not need to be the original pretraining data. Good options are open instruction datasets with a suitable licence, or answers the base model itself wrote for varied prompts — training on those tells the model "keep behaving as you did".
Measure it
Keep a small regression suite and run it on the base and the fine-tuned model with the same harness: a few hundred items from general knowledge, maths, instruction following and safety refusals, plus real general questions from your own traffic. Section 3 covers how to recover once forgetting has happened.
A real-life example
An e-commerce team fully fine-tunes an 8B model for Hindi support: 3 epochs, learning rate 2e-5, 12,000 chats. Support quality improves. A week later, agents notice the model now gets EMI calculations wrong and answers English questions in Hindi.
They run their 400-item general suite: the score dropped from 78 to 61 (their numbers). They retrain with a rank-16 LoRA, 2 epochs, and 10% replay data — Hindi and English general instructions answered by the original instruct model. The support score stays almost the same, and the general suite comes back to 76. The general suite now runs automatically after every training job.
Follow-up questions to expect
- "Does LoRA remove forgetting completely?" — No. It forgets less, and you can switch the adapter off to get the base model back, but the adapted model can still lose skills. You still measure.
- "What is EWC?" — Elastic Weight Consolidation adds a penalty for changing weights that were important for old tasks. It is a classic continual-learning method but is rarely used for LLMs, where replay and PEFT are simpler.
- "How much replay data?" — Start around 10% and tune: too little and forgetting returns; too much and the task gains shrink.