Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is catastrophic forgetting, and what techniques help avoid it while fine-tuning?


What you need to know

Why it happens

All skills share the same weights. During fine-tuning, the gradient only says "get better at this data". It has no information about maths, coding or politeness, so the update can move weights those skills depended on. The more the weights move — high learning rate, many epochs, full fine-tuning — the more is lost.

What gets forgotten

  • General instruction following — ignoring formatting requests, answering in the wrong language.
  • Reasoning and maths — especially when the new data has none.
  • Safety behaviour. Research in 2023 (Qi et al.) showed that fine-tuning an aligned model, even on harmless data, can weaken its refusals. Re-test safety after every fine-tune.
  • Other languages — a model tuned only on English support chats may get worse at Hindi.

Techniques that help

TechniqueHow it helpsCost
LoRA / PEFTBase weights frozen; the update is small and can be switched offLow
Low learning rate, 1–3 epochsLess weight movementFree
Replay (5–20% general data)The loss keeps rewarding old skillsSome extra training time
KL or distillation term against the basePenalises drifting from the base model's outputs on general promptsNeeds the base model in the loop
Weight averaging / lower adapter scalePulls the model part-way back toward the baseSmall loss of task gain
Keep facts in retrievalLess new knowledge pushed into weightsA retrieval system

Replay data does not need to be the original pretraining data. Good options are open instruction datasets with a suitable licence, or answers the base model itself wrote for varied prompts — training on those tells the model "keep behaving as you did".

Measure it

Keep a small regression suite and run it on the base and the fine-tuned model with the same harness: a few hundred items from general knowledge, maths, instruction following and safety refusals, plus real general questions from your own traffic. Section 3 covers how to recover once forgetting has happened.

A real-life example

An e-commerce team fully fine-tunes an 8B model for Hindi support: 3 epochs, learning rate 2e-5, 12,000 chats. Support quality improves. A week later, agents notice the model now gets EMI calculations wrong and answers English questions in Hindi.

They run their 400-item general suite: the score dropped from 78 to 61 (their numbers). They retrain with a rank-16 LoRA, 2 epochs, and 10% replay data — Hindi and English general instructions answered by the original instruct model. The support score stays almost the same, and the general suite comes back to 76. The general suite now runs automatically after every training job.

Follow-up questions to expect

  • "Does LoRA remove forgetting completely?" — No. It forgets less, and you can switch the adapter off to get the base model back, but the adapted model can still lose skills. You still measure.
  • "What is EWC?" — Elastic Weight Consolidation adds a penalty for changing weights that were important for old tasks. It is a classic continual-learning method but is rarely used for LLMs, where replay and PEFT are simpler.
  • "How much replay data?" — Start around 10% and tune: too little and forgetting returns; too much and the task gains shrink.