Course Content
Transfer Learning and Pretraining
3 sections · 7 lessons
What Is Transfer Learning, and Why Use It?
A small team is building a skin-lesion classifier. They have 1,800 dermatoscopy images, seven diagnostic classes, and a ResNet-50. They train it from scratch for 60 epochs. The training curve looks wonderful — 99.4% accuracy on the training set. Then they check the validation set: 61%. Always predicting the most common class would give roughly 40% on a dataset this imbalanced, so the model has learned almost nothing generalisable. It has memorised 1,800 photographs.
They try the usual remedies. More dropout: validation creeps to 64%. Heavier augmentation: 66%. A smaller network: 68%, but now training accuracy has fallen too, so they are simply losing on both sides. Weeks disappear.
Then someone changes four lines. Instead of random initial weights, they load weights from a ResNet-50 that was trained on ImageNet — 1.28 million photographs of dogs, cars, mushrooms and washing machines, none of them medical. They freeze everything except the final classification layer and train for twelve minutes on a single GPU.
Validation accuracy: 87%.
Nothing about the dataset changed. No new images were collected. The architecture is identical. The only difference is where the weights started. That gap — 61% to 87% — is what this topic is about, and the interesting question is not that it works but why a network that has never seen a mole should be so much better at diagnosing one than a network trained specifically on moles.
Why training from scratch fails on 1,800 images
Start with arithmetic, because the arithmetic is brutal and it explains everything that follows.
ResNet-50 has about 25.6 million trainable parameters. The team's dataset has 1,800 images. Each image, at 224×224 with three colour channels, is 150,528 numbers, but from the optimiser's point of view what matters is how many independent constraints the data places on those 25.6 million knobs. With 1,800 labelled examples and seven classes, the labels alone carry at most a few thousand bits of information.
Divide it out: 25,600,000 parameters ÷ 1,800 examples ≈ 14,200 free parameters per training example. A system with fourteen thousand degrees of freedom per constraint is not solving an equation; it is choosing arbitrarily from an enormous space of solutions that all fit the training data perfectly. Almost all of those solutions are memorisation. Only a vanishingly small fraction happen to generalise.
Overfitting is not a bug that appears when you train too long. It is the default outcome whenever the model has far more capacity than the data has constraints — training longer just lets it finish the job.
The classical fixes all work by removing capacity or adding constraints: a smaller network, stronger weight decay, dropout, augmentation. Every one of them buys generalisation by paying in expressiveness. That is why the team's numbers crept up in small increments and then stalled — they were trading one problem for another.
Transfer learning attacks the same imbalance from the opposite side. Instead of shrinking the model, it pre-constrains it. The weights arrive already shaped by 1.28 million images, so the 1,800 medical images no longer have to determine 25.6 million numbers from nothing. They only have to adjust a model that is already in roughly the right place.
What transfer learning actually is
Now the definition, which will land properly because you have felt the problem it solves.
Transfer learning is the reuse of knowledge learned while solving one problem to improve learning on a different but related problem. In practice: you take the internal representations a model built on a large dataset, and you use them as the starting point for a smaller, different task.
To reason about it precisely you need two pieces of vocabulary, and they are worth separating carefully because most confusion about transfer learning comes from mixing them up.
Domain and task are different things
A domain is the data and its distribution. Formally D={X,P(X)}: the space of possible inputs, plus how likely each input is. Photographs taken with a phone camera in daylight are one domain. Dermatoscope images under a polarised light source are a different domain, even though both are "images of skin".
A task is what you are predicting. Formally T={Y,f}: the label space plus the function mapping inputs to labels. "Which of 1,000 everyday objects is this?" is one task. "Which of seven lesion types is this?" is a different task, even on identical images.
So you always have four objects in play: a source domain, a source task, a target domain, a target task. Transfer learning is what you do when the source pair and the target pair are not the same, and the difference between them is the transfer gap.
| Situation | Domain | Task | Name | Example |
|---|---|---|---|---|
| 1 | Same | Same | Ordinary supervised learning | Train and deploy on the same data distribution |
| 2 | Different | Same | Domain adaptation | Pedestrian detector trained on daylight footage, deployed at night |
| 3 | Same | Different | Task transfer | ImageNet classifier reused as an object detector on natural photos |
| 4 | Different | Different | Full transfer | ImageNet photos → dermatoscopy lesion grading |
The skin-lesion team is in situation 4, the hardest one — and it still gave them 26 percentage points. That is the headline result of the whole field.
Why it works: features are layered, and the early layers are generic
The deepest reason transfer learning works is that a convolutional network does not learn one monolithic "cat detector". It learns a hierarchy, and the bottom of that hierarchy is not about cats at all.
If you visualise the filters in the first convolutional layer of almost any vision network — trained on ImageNet, on faces, on satellite imagery, on X-rays — you find the same things: oriented edge detectors and colour-opponent blobs. They look like Gabor filters. Nobody designed them; every network rediscovers them, because they are the optimal first step for almost any visual task.
| Depth | What the layer responds to | How task-specific | Transfers to a new task? |
|---|---|---|---|
| Layer 1–2 | Edges at various orientations, colour blobs | Essentially universal | Almost always, unchanged |
| Layer 3–4 | Corners, junctions, simple textures | Near-universal for images | Almost always, unchanged |
| Middle blocks | Object parts — wheels, eyes, fur patches, mesh patterns | Somewhat domain-specific | Usually, with light adjustment |
| Late blocks | Whole-object configurations | Strongly source-specific | Often needs fine-tuning |
| Classifier head | The 1,000 ImageNet categories | Entirely source-specific | Never — always replaced |
The landmark study here is Yosinski and colleagues in 2014, who split ImageNet into two halves and measured transferability layer by layer. Three findings from that work are worth carrying with you:
- Generality decays smoothly with depth. There is no clean boundary between "generic" and "specific" layers — the transition is gradual, which is why "how many layers should I freeze?" has no universal answer.
- Freezing mid-level layers can hurt more than freezing early or late ones. Adjacent layers develop co-adapted features — neuron A only makes sense given what neuron B is doing. Cutting a frozen/trainable boundary through the middle of a co-adapted group breaks the collaboration, and the trainable side cannot repair it because the frozen side will not move.
- Transferred initialisation helps even after full fine-tuning. Even when every weight is eventually allowed to change, starting from pretrained weights beats starting from random ones. The benefit is not only the frozen features; it is the starting point.
Pretraining does not give you a model that knows about your problem. It gives you a model that already knows what images look like — and that turns out to be most of the work.
Why it works: the statistics and the optimisation
Two more mechanisms are worth spelling out, because they explain when transfer will help and when it will not.
The generalisation bound tightens
Classical learning theory says that the gap between training error and true error grows roughly with d/n, where d is the effective capacity of the hypothesis space and n is the number of training examples. You cannot change n — collecting more dermatoscopy images is expensive. But freezing the backbone changes d dramatically.
Work it through. Full ResNet-50: d≈25,600,000. Freeze everything and train only a new head mapping the 2048-dimensional feature vector to 7 classes: d=2048×7+7=14,343 trainable parameters. The ratio is about 1,785, and since the bound scales with the square root, the generalisation gap tightens by roughly 1785≈42×.
That is a hand-wavy bound and real networks beat it, but the direction is exactly right and it matches what the team observed: their frozen-backbone model showed a train/validation gap of a few points instead of thirty-eight.
The optimisation problem gets easier
Random initialisation drops you at a random point in a 25-million-dimensional loss landscape. Pretrained weights drop you in a basin that already produces useful representations. Three consequences follow:
- Fewer steps to convergence. Training ResNet-50 on ImageNet from scratch takes roughly 90 epochs over 1.28M images — about 115 million image presentations. Fine-tuning on 1,800 images for 10 epochs is 18,000 presentations. That is a factor of about 6,400 less compute.
- Less sensitivity to hyperparameters. From-scratch training on small data is notoriously fragile — the wrong learning rate diverges, the wrong initialisation stalls. Fine-tuning tolerates a much wider range.
- Better-conditioned gradients. A network with sensible features produces informative gradients from the first batch, instead of spending thousands of steps learning that edges exist.
From scratch versus transfer: the numbers side by side
Here is the same seven-class, 1,800-image problem run three ways on one GPU. These are typical figures for this kind of dataset, not a single lucky run.
| Approach | Trainable params | Epochs | Wall-clock | Train acc | Val acc | Train − Val gap |
|---|---|---|---|---|---|---|
| ResNet-50 from scratch | 25.6M | 60 | ~70 min | 99.4% | 61% | 38.4 pts |
| Frozen backbone + new head | 14.3K | 15 | ~12 min | 90.1% | 87.0% | 3.1 pts |
| Pretrained + full fine-tune (low LR) | 25.6M | 20 | ~28 min | 97.6% | 91.4% | 6.2 pts |
Read the third column against the last one. The from-scratch model trained four times longer and ended up with a gap twelve times wider. Notice also that full fine-tuning beats the frozen head here — with 1,800 images that is a close call, and on a dataset of 300 images the frozen version would usually win. The dataset size decides.
When transfer learning helps, and when it does not
Transfer learning is not free and it is not universal. The honest version of the decision looks like this.
| Condition | Transfer is a strong win | Transfer is marginal or harmful |
|---|---|---|
| Target dataset size | Under ~50k labelled examples | Millions of examples — from scratch catches up |
| Input modality | Same as source (RGB images, natural text) | Radically different (raw radar returns, 12-lead ECG, point clouds) |
| Semantic distance | Source features are plausibly relevant | Source and target share no visual or linguistic structure |
| Compute budget | Limited — you cannot afford 90 ImageNet epochs | Effectively unlimited |
| Input resolution | Near the pretraining resolution (224×224) | Wildly different, e.g. 4000×4000 histopathology slides |
| Time to first result | You need a working baseline this week | Research setting with months available |
The "millions of examples" row deserves a caveat that people frequently get wrong: even when you have enough data that from-scratch training eventually matches transfer on final accuracy, transfer usually still converges faster. The benefit shifts from accuracy to compute cost, and compute cost is still worth having.
The modality row is the one that genuinely blocks transfer. ImageNet features are built for photographs of three-dimensional objects lit from above. A 12-lead ECG trace is a one-dimensional signal in which the informative structure is millisecond-scale timing. Edge detectors have nothing to say about it. This is where people force ImageNet weights onto spectrograms or sensor traces, get 2% improvement, and conclude transfer learning is overrated — the technique was fine, the source was wrong.
The vocabulary, precisely
| Term | What it means | The distinction that matters |
|---|---|---|
| Pretrained model | A network whose weights come from training on a large source dataset | Weights, not architecture. "ResNet-50" is an architecture; "ResNet-50 with ImageNet weights" is a pretrained model. |
| Backbone | The feature-producing body of the network, with its original classifier removed | What you keep. The head is what you throw away. |
| Head | The final layer(s) mapping features to your labels | Always replaced, because your label space differs from the source's. |
| Feature extraction | Freeze the backbone; train only a new head | Backbone weights never change. Cheap, low-variance, caps out lower. |
| Fine-tuning | Allow some or all backbone weights to update, usually at a small learning rate | Backbone weights change. Higher ceiling, higher overfitting risk. |
| Domain shift | Psource(X)=Ptarget(X) — the input distributions differ | About the inputs, not the labels. Can exist even when the task is identical. |
| Negative transfer | Pretrained initialisation makes the target model worse than random init | Rare but real. Always worth a from-scratch control run. |
Where this has already changed the field
The ImageNet effect
After 2012, "load ImageNet weights" stopped being a technique and became the default. The reason is that ImageNet-1k is unusually well-suited as a source: 1.28 million images, 1,000 fine-grained categories that force the network to distinguish 120 dog breeds from each other, and enough visual variety that the learned features are genuinely general. Object detection, segmentation, pose estimation, medical imaging and satellite analysis all standardised on ImageNet-pretrained backbones. Detection frameworks are built around a swappable pretrained backbone as a first-class design element.
The language model effect
Text went through the same transition, later and harder. Before 2018, an NLP system for a new task typically meant task-specific architecture, task-specific features and tens of thousands of labelled examples. BERT changed the economics: pretrain one encoder on billions of words of unlabelled text by predicting masked words, then fine-tune it on a few thousand labelled examples per task. Tasks that had resisted progress for years moved several points overnight, and the reason was purely that the model arrived already knowing syntax, coreference and word sense.
The important structural difference is the source labels. ImageNet pretraining needs 1.28 million human annotations. Language pretraining needs none — the text supervises itself, because the next word (or the masked word) is already there. That is why language models scaled to trillions of tokens while image pretraining datasets stayed in the millions, and it is why self-supervised pretraining has since become the dominant approach for images too.
Multimodal
Models trained to match images with their captions learn a shared representation space for both. The practical payoff is zero-shot classification: you can classify images into categories the model was never explicitly trained on, by comparing the image's representation against the representation of the text "a photo of a <category>". This is transfer learning with the fine-tuning step removed entirely.
Three ways it goes wrong
These are the failure modes worth recognising by symptom, because each has a different fix and applying the wrong one wastes days.
| Failure | Symptom you observe | Underlying cause | What actually fixes it |
|---|---|---|---|
| Negative transfer | Pretrained model performs worse than a from-scratch control, or plateaus early and low | Source features are irrelevant or actively misleading for the target | Change source (domain-specific pretraining), or train from scratch, or self-supervised pretrain on your own unlabelled data |
| Catastrophic forgetting | Loss spikes in the first few hundred steps; final accuracy below the frozen-backbone baseline | Learning rate too high — large gradients destroy pretrained features before the random head has stabilised | Lower LR by 10–100×, warm up, freeze the backbone for the first epoch while the head settles, use layer-wise decaying learning rates |
| Domain shift | Training and validation both look fine; production accuracy is far lower | Deployment inputs differ from training inputs — new scanner, new lighting, new camera | Domain adaptation, recalibrating normalisation statistics on target data, augmentation that simulates the shift, collecting target-domain data |
Catastrophic forgetting is the one that catches almost everyone the first time. The mechanism is worth understanding rather than memorising: your new head is randomly initialised, so its first predictions are nonsense, so the loss is large, so the gradients flowing back into the carefully-trained backbone are large. At a learning rate suitable for from-scratch training — say 0.1 with SGD — those gradients scramble features that took 115 million image presentations to build. The damage is done in the first fifty steps, before you have looked at a single validation number.
What this means when you start a project
The practical consequence is that "should I use transfer learning?" is almost never the right question. The right question is "which source, and how much of it do I let move?" — and you answer it in a fixed order.
- Establish the frozen baseline first. Load a pretrained backbone, replace the head, freeze everything else, train for ten to fifteen epochs. This takes minutes and gives you a number that everything else must beat. Teams that skip this step have no idea whether their elaborate fine-tuning schedule is helping.
- Run a from-scratch control if the domain is unusual. Medical, scientific, sensor or audio data all deserve this check. If from-scratch matches or beats the pretrained model, you have found negative transfer and you should change source, not tune harder.
- Choose your source by modality first, semantics second. An X-ray-pretrained backbone beats an ImageNet-pretrained one for radiology, and both beat an ImageNet backbone applied to raw waveform audio. Match the input type before you worry about label similarity.
- Only then unfreeze, and do it at a learning rate one or two orders of magnitude below what you would use from scratch. Watch the first epoch closely; a loss spike means you have already gone too far.
- Budget your labelling effort against the curve, not against a target. Going from 500 to 2,000 images typically buys far more than going from 20,000 to 80,000, because the pretrained features are already carrying the early part of the curve. Measure your accuracy at 25%, 50% and 100% of your current data before commissioning more annotation — the slope of that curve tells you whether more labels are worth the money.
The team with the skin lesions eventually reached 91.4% by unfreezing the last two residual blocks at a learning rate of 1e-4 while keeping the earlier layers frozen. They never collected a single additional image. Everything they gained came from making better use of a network that had spent its life looking at photographs of dogs.