LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How can catastrophic forgetting be mitigated in LLMs?


What you need to know

Why it happens

Fine-tuning data is narrow: 5,000 examples of one task, one style, one language. Every gradient step pushes weights towards that narrow pattern, and nothing pushes back to protect other skills. Many epochs and high learning rates make it worse.

Mitigations, in the order I would try them

  1. Train gently — LoRA instead of full fine-tuning; learning rate around 1e-4 for LoRA (lower, around 1e-5, for full fine-tuning); 1–3 epochs; early stopping.
  2. Mix data (rehearsal) — include a share of general instruction data, such as 10–30% of each batch, so the model keeps practising old behaviours.
  3. Regularise towards the original — L2 penalty towards the starting weights, Elastic Weight Consolidation (penalises changes to weights important for earlier tasks), or a KL penalty to the original model's outputs.
  4. Freeze more — freeze the lower layers or the embeddings.
  5. Merge weights — interpolate between the original and fine-tuned weights (for example 70% fine-tuned, 30% original) to trade task gain for retained skill.
  6. Evaluate broadly — a regression suite of general tasks, safety tests and other languages, run before and after.

Question the need to fine-tune

If the goal is new knowledge, retrieval avoids forgetting entirely because the weights do not change. Fine-tune for behaviour, format or style.

A real-life example

An e-commerce company fully fine-tunes an open model for 5 epochs on 20,000 product Q&A pairs, all in English. Product-question accuracy rises. But afterwards, the search assistant answers Hindi questions in English, ignores "reply in JSON" instructions from the app, and refuses less reliably — none of which appeared in the training data, so the task evaluation never noticed.

The retrain uses LoRA (r = 16), 2 epochs, a learning rate of 1e-4, and batches that are 80% product Q&A and 20% general instruction data including Hindi and JSON-format examples. They also add a 400-case regression suite (Hindi, JSON, safety, general questions). The new model keeps almost all of the product gain, and the regression suite now matches the base model within a point.

Follow-up questions to expect

  • "How do you detect forgetting?" — Run the same broad evaluation set on the base and fine-tuned models and compare per category, not just the average.
  • "What is Elastic Weight Consolidation?" — A penalty that makes it expensive to move weights that were important for earlier tasks, with importance estimated from the Fisher information.
  • "Does instruction tuning itself cause forgetting?" — It can reduce some raw pretrained abilities (sometimes called the alignment tax), which is why labs mix pretraining data or use KL penalties during those stages.