Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

If domain fine-tuning hurts general capability, how do you recover from catastrophic forgetting?


What you need to know

Step 1: measure

Run knowledge, maths, instruction-following and safety checks on both models with the same prompts and settings. Note which skills dropped — that decides the fix. Section 2 covers why forgetting happens.

Step 2: no-retraining fixes

Weight interpolation (called WiSE-FT, 2022) mixes the two models:

Text
W = (1 - a) * W_base + a * W_finetuned        a between 0 and 1

For LoRA this is simply scaling the adapter, because W_finetuned = W_base + update. PEFT 0.21 has a helper:

Python
from peft.helpers import rescale_adapter_scalewith rescale_adapter_scale(model, multiplier=0.6):    run_eval(model, general_suite)   # adapter at 60% strength    run_eval(model, domain_suite)

Try a few values (0.4, 0.6, 0.8) and pick the one with the best balance. This often recovers much of the general skill for a small task loss.

Step 3: retrain more gently

ChangeEffect
Add 10–30% replay dataOld skills stay in the loss
Lower learning rate, fewer epochsLess movement from the base
Lower rank, or target fewer modulesLess capacity to overwrite
KL or distillation term against the base on general promptsDirectly penalises drift

Replay can be answers the base model wrote to varied prompts; training on them says "keep doing what you did".

Step 4: route instead of merging

Keep the adapter separate. A small classifier or rule sends domain questions through the adapter and general questions to the plain base model. Zero forgetting on general queries, at the cost of a routing step that can misroute.

A real-life example

The Mumbai law firm fully fine-tunes its assistant on 20,000 legal tasks. Clause accuracy is up, but the firm's general suite shows email drafting and Hindi translation have dropped sharply.

They first interpolate the fine-tuned weights with the base at a = 0.6: most of the general skill returns, but clause accuracy falls 2 points. They then retrain as a LoRA adapter with 15% replay data drawn from the base model's own answers to general drafting and translation prompts. Clause accuracy matches the full fine-tune, and the general suite is within 1 point of the base. The general suite now runs in CI, and a drop of more than 2 points blocks release.

Follow-up questions to expect

  • "Why does interpolation work at all?" — The base and fine-tuned models usually sit in the same low-loss region, so points between them stay good at both tasks. It fails if the fine-tune moved very far.
  • "How do you choose the replay share?" — Start around 10–15%, and raise it until general scores recover without the domain score dropping too much.
  • "Can you fix it with prompting?" — Sometimes partly, with a system prompt restoring the lost behaviour, but that treats the symptom.