Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is transfer learning, why is it needed, and which pretrained models are commonly used?
What you need to know
Why pretrained features transfer
A CNN trained on ImageNet, over a million photos in 1,000 classes, learns a hierarchy. Its first layers detect edges and colour blobs; middle layers detect textures and simple shapes; later layers detect object parts such as wheels, eyes and fabric patterns. Only the last layer is specific to ImageNet's 1,000 classes. The early and middle features are useful for almost any photo task, from product images to crop disease.
Language models work the same way. A model pretrained to predict words on huge amounts of text has learned grammar, word meanings and a lot of world knowledge, which any text task can build on.
The two main ways to use it
- Feature extraction: freeze the pretrained layers and train only a new output head. Fast, needs little data, and hard to overfit.
- Fine-tuning: also update some or all pretrained layers, with a small learning rate. Usually more accurate when you have more data or your images look different from the pretraining data. The next lesson covers it.
1import torch.nn as nn2from torchvision.models import resnet50, ResNet50_Weights34weights = ResNet50_Weights.DEFAULT5model = resnet50(weights=weights)6for p in model.parameters():7 p.requires_grad = False # freeze the backbone8model.fc = nn.Linear(model.fc.in_features, 12) # new head: 12 classes, trainable9preprocess = weights.transforms() # the resize and normalisation it expectsweights= is the current torchvision API; the old pretrained=True argument is deprecated. The new nn.Linear is created after the freeze loop, so its parameters are trainable by default. weights.transforms() returns the exact preprocessing the model was trained with; using different normalisation is a common silent bug.
Commonly used pretrained models
| Domain | Models |
|---|---|
| Image classification and features | ResNet, EfficientNet, ConvNeXt, ViT, DINOv2 |
| Object detection and segmentation | YOLO family, DETR-style models, Mask R-CNN, SAM |
| Image and text together | CLIP, SigLIP |
| Text understanding | BERT, RoBERTa, DeBERTa, sentence-transformers for embeddings |
| Text generation | Open LLMs such as Llama, Mistral, Qwen and Gemma |
| Speech | Whisper, wav2vec 2.0 |
Libraries such as timm (image models) and Hugging Face transformers make most of these one line to load.
When transfer learning helps less
- When your data is very different from the pretraining data, such as radar or multispectral satellite bands. It usually still helps a bit, but less.
- When you have a huge labelled dataset and the compute to use it; the pretrained start matters less.
- When the pretrained model is far bigger than your latency budget allows. Then you distil it into a smaller model or pick a smaller backbone.
A real-life example
A handloom brand selling online has 1,800 product photos across 12 categories, such as sarees, stoles, cushion covers and table runners. Training a ResNet-50 from random weights on these images gives about 55% validation accuracy after hours of training: 25 million parameters and 1,800 images is a recipe for memorising.
Using the snippet above, the team freezes an ImageNet-pretrained ResNet-50 and trains only the new 12-class head. After 10 minutes on one GPU, validation accuracy is 88%. The pretrained filters already recognise weaves, fringes, borders and folds; the head just learns which combinations mean "stole" rather than "table runner". The next step, fine-tuning the last block, is where the remaining gains come from.
Follow-up questions to expect
- "Which layers transfer best?" — The early ones, because they learn general patterns. The later layers are more specific to the original task, so they are the first to unfreeze and retrain.
- "What is negative transfer?" — When starting from pretrained weights makes results worse than training from scratch, usually because the source and target data are very different. It is rare with large, diverse pretraining.
- "Is using a pretrained model's embeddings transfer learning?" — Yes. Using a frozen model to produce features, such as sentence embeddings for search or clustering, is the simplest form.