Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
If domain fine-tuning hurts general capability, how do you recover from catastrophic forgetting?
What you need to know
Step 1: measure
Run knowledge, maths, instruction-following and safety checks on both models with the same prompts and settings. Note which skills dropped — that decides the fix. Section 2 covers why forgetting happens.
Step 2: no-retraining fixes
Weight interpolation (called WiSE-FT, 2022) mixes the two models:
W = (1 - a) * W_base + a * W_finetuned a between 0 and 1For LoRA this is simply scaling the adapter, because W_finetuned = W_base + update. PEFT 0.21 has a helper:
1from peft.helpers import rescale_adapter_scale23with rescale_adapter_scale(model, multiplier=0.6):4 run_eval(model, general_suite) # adapter at 60% strength5 run_eval(model, domain_suite)Try a few values (0.4, 0.6, 0.8) and pick the one with the best balance. This often recovers much of the general skill for a small task loss.
Step 3: retrain more gently
| Change | Effect |
|---|---|
| Add 10–30% replay data | Old skills stay in the loss |
| Lower learning rate, fewer epochs | Less movement from the base |
| Lower rank, or target fewer modules | Less capacity to overwrite |
| KL or distillation term against the base on general prompts | Directly penalises drift |
Replay can be answers the base model wrote to varied prompts; training on them says "keep doing what you did".
Step 4: route instead of merging
Keep the adapter separate. A small classifier or rule sends domain questions through the adapter and general questions to the plain base model. Zero forgetting on general queries, at the cost of a routing step that can misroute.
A real-life example
The Mumbai law firm fully fine-tunes its assistant on 20,000 legal tasks. Clause accuracy is up, but the firm's general suite shows email drafting and Hindi translation have dropped sharply.
They first interpolate the fine-tuned weights with the base at a = 0.6: most of the general skill returns, but clause accuracy falls 2 points. They then retrain as a LoRA adapter with 15% replay data drawn from the base model's own answers to general drafting and translation prompts. Clause accuracy matches the full fine-tune, and the general suite is within 1 point of the base. The general suite now runs in CI, and a drop of more than 2 points blocks release.
Follow-up questions to expect
- "Why does interpolation work at all?" — The base and fine-tuned models usually sit in the same low-loss region, so points between them stay good at both tasks. It fails if the fine-tune moved very far.
- "How do you choose the replay share?" — Start around 10–15%, and raise it until general scores recover without the domain score dropping too much.
- "Can you fix it with prompting?" — Sometimes partly, with a system prompt restoring the lost behaviour, but that treats the symptom.