Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is fine-tuning, and why is it used in deep learning?
What you need to know
Fine-tuning versus feature extraction
With feature extraction, the pretrained layers are frozen and only the head learns. With fine-tuning, gradients also flow into the pretrained layers, so their features adjust to your data. Fine-tuning is usually more accurate when your images or text look different from the pretraining data, as long as you have enough examples to support the extra freedom.
A common two-stage recipe
- Train the head — freeze the backbone and train only the new head for a few epochs. The head starts random, and its large early gradients would otherwise damage the pretrained layers.
- Unfreeze the top — unfreeze the last block or two, which hold the most task-specific features.
- Use lower learning rates for older layers — for example 1e-4 for the unfrozen block and 1e-3 for the head. This is called discriminative learning rates.
- Stop early — watch validation loss; fine-tuning can overfit quickly on small datasets.
1for p in model.layer4.parameters():2 p.requires_grad = True # unfreeze the last ResNet block34optimizer = torch.optim.AdamW([5 {"params": model.layer4.parameters(), "lr": 1e-4},6 {"params": model.fc.parameters(), "lr": 1e-3},7], weight_decay=0.01)Each parameter group gets its own learning rate. Layers 1 to 3 stay frozen and are not passed to the optimiser.
Catastrophic forgetting
Small learning rates, few epochs, freezing early layers and parameter-efficient methods all reduce it.
Fine-tuning large language models
Updating every weight of a multi-billion-parameter model needs a lot of GPU memory, because optimiser state is stored for every parameter. Parameter-efficient fine-tuning (PEFT) freezes the base model and trains a small number of new weights. LoRA adds small low-rank matrices next to existing weight matrices; QLoRA does the same on a base model loaded in 4-bit precision to save more memory. Often well under 1% of parameters are trained, and the adapter file is small enough to swap per customer or task.
Fine-tuning is good at teaching behaviour: a format, a tone, a labelling scheme, domain vocabulary. It is a poor way to add facts that change, such as today's prices or policies, because you would have to retrain every time. For those, use retrieval-augmented generation (RAG).
A real-life example
A diagnostics chain fine-tunes an ImageNet-pretrained ResNet-50 to triage 20,000 labelled chest X-rays as "urgent" or "routine". X-rays are grayscale and look very different from ImageNet photos, so feature extraction alone reaches only 84% validation accuracy: the frozen features were built for dogs and cars.
Stage 2 unfreezes layer4 with learning rate 1e-4 and the head at 1e-3, and accuracy rises to 90%. Stage 3 also unfreezes layer3 at 3e-5 and reaches 91.5%, with early stopping at epoch 9. An earlier attempt that unfroze everything at 1e-3 from the start scored 86%: the large, random-head gradients wrecked the pretrained features in the first epoch. Same model, same data; the learning rates and order made the difference.
Follow-up questions to expect
- "How much data do you need to fine-tune?" — It depends on how different your data is from the pretraining data. A few hundred examples per class can be enough for head-only training on similar images; fully fine-tuning a large model on very different data needs much more.
- "What is the difference between LoRA and full fine-tuning?" — Full fine-tuning updates every weight. LoRA freezes them and learns small low-rank updates added to chosen layers, which uses far less memory and gives a small adapter file, usually with similar quality on narrow tasks.
- "Should batch norm layers be frozen when fine-tuning on small batches?" — Often yes. Keeping them in eval mode preserves the pretrained running statistics, which are more reliable than statistics from small batches of your data.