Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is fine-tuning, and why is it used in deep learning?


What you need to know

Fine-tuning versus feature extraction

With feature extraction, the pretrained layers are frozen and only the head learns. With fine-tuning, gradients also flow into the pretrained layers, so their features adjust to your data. Fine-tuning is usually more accurate when your images or text look different from the pretraining data, as long as you have enough examples to support the extra freedom.

A common two-stage recipe

  1. Train the head — freeze the backbone and train only the new head for a few epochs. The head starts random, and its large early gradients would otherwise damage the pretrained layers.
  2. Unfreeze the top — unfreeze the last block or two, which hold the most task-specific features.
  3. Use lower learning rates for older layers — for example 1e-4 for the unfrozen block and 1e-3 for the head. This is called discriminative learning rates.
  4. Stop early — watch validation loss; fine-tuning can overfit quickly on small datasets.
Python
for p in model.layer4.parameters():    p.requires_grad = True                     # unfreeze the last ResNet blockoptimizer = torch.optim.AdamW([    {"params": model.layer4.parameters(), "lr": 1e-4},    {"params": model.fc.parameters(),     "lr": 1e-3},], weight_decay=0.01)

Each parameter group gets its own learning rate. Layers 1 to 3 stay frozen and are not passed to the optimiser.

Catastrophic forgetting

Small learning rates, few epochs, freezing early layers and parameter-efficient methods all reduce it.

Fine-tuning large language models

Updating every weight of a multi-billion-parameter model needs a lot of GPU memory, because optimiser state is stored for every parameter. Parameter-efficient fine-tuning (PEFT) freezes the base model and trains a small number of new weights. LoRA adds small low-rank matrices next to existing weight matrices; QLoRA does the same on a base model loaded in 4-bit precision to save more memory. Often well under 1% of parameters are trained, and the adapter file is small enough to swap per customer or task.

Fine-tuning is good at teaching behaviour: a format, a tone, a labelling scheme, domain vocabulary. It is a poor way to add facts that change, such as today's prices or policies, because you would have to retrain every time. For those, use retrieval-augmented generation (RAG).

A real-life example

A diagnostics chain fine-tunes an ImageNet-pretrained ResNet-50 to triage 20,000 labelled chest X-rays as "urgent" or "routine". X-rays are grayscale and look very different from ImageNet photos, so feature extraction alone reaches only 84% validation accuracy: the frozen features were built for dogs and cars.

Stage 2 unfreezes layer4 with learning rate 1e-4 and the head at 1e-3, and accuracy rises to 90%. Stage 3 also unfreezes layer3 at 3e-5 and reaches 91.5%, with early stopping at epoch 9. An earlier attempt that unfroze everything at 1e-3 from the start scored 86%: the large, random-head gradients wrecked the pretrained features in the first epoch. Same model, same data; the learning rates and order made the difference.

Follow-up questions to expect

  • "How much data do you need to fine-tune?" — It depends on how different your data is from the pretraining data. A few hundred examples per class can be enough for head-only training on similar images; fully fine-tuning a large model on very different data needs much more.
  • "What is the difference between LoRA and full fine-tuning?" — Full fine-tuning updates every weight. LoRA freezes them and learns small low-rank updates added to chosen layers, which uses far less memory and gives a small adapter file, usually with similar quality on narrow tasks.
  • "Should batch norm layers be frozen when fine-tuning on small batches?" — Often yes. Keeping them in eval mode preserves the pretrained running statistics, which are more reliable than statistics from small batches of your data.