Course Content
Deep Learning Essentials
13 sections · 61 lessons
Why are deep neural networks more prone to overfitting?
What you need to know
Generalisation is the opposite: doing well on data the model has never seen. The only honest measure of it is a validation set (used to tune) and a test set (touched once at the end), both kept apart from training.
Why deep networks in particular
- Parameters far outnumber examples. A ResNet-18 has about 11 million parameters. If you train it from scratch on 2,000 photos, it has thousands of parameters per photo. That is enough to store an answer for every image.
- They can fit anything, even noise. Researchers showed that standard image networks can reach 100% training accuracy on images whose labels were randomly shuffled. The network memorised pure noise. So a perfect training score proves nothing about learning.
- Label noise. If 5% of training labels are wrong, a big network will eventually learn those wrong labels too.
- Training too long. Early epochs learn broad patterns that most examples share. Later epochs fit the rare, odd examples one by one. That later phase is where overfitting happens.
- High-dimensional input, few samples. A 224×224 image has 150,528 values. With few examples, there are many accidental patterns (a watermark, a background colour) that happen to predict the label.
How to see it
Plot training and validation loss for each epoch:
epoch train_loss val_loss 1 1.90 1.85 5 0.60 0.72 10 0.20 0.65 <- best validation loss 20 0.03 0.95 <- overfitting: gap is growing 40 0.001 1.40The gap between the two curves is the signal. Both high means underfitting. Training low and validation rising means overfitting.
The toolkit, in rough order of value
- More real data, or cleaner labels.
- Data augmentation — flips, crops, colour changes, noise, so each epoch sees a slightly different image.
- Transfer learning — start from a pretrained model, so far fewer parameters must be learned from your data.
- Early stopping — keep the weights from the epoch with the best validation loss.
- Weight decay (L2) and dropout — limit how much the network can rely on large weights or single neurons.
- A smaller model, if the above are not enough.
A minimal early-stopping loop looks like this:
1best, bad_epochs, patience = float("inf"), 0, 52for epoch in range(100):3 train_one_epoch(model, train_loader)4 val_loss = evaluate(model, val_loader)5 if val_loss < best:6 best, bad_epochs = val_loss, 07 torch.save(model.state_dict(), "best.pt")8 else:9 bad_epochs += 110 if bad_epochs == patience:11 breaktrain_one_epoch and evaluate are your own functions. The loop stops after 5 epochs with no improvement and keeps the best checkpoint, not the last one.
A real-life example
A team builds a crop-disease classifier with 1,200 photos of cotton leaves, all taken at one research farm. Training from scratch, the CNN reaches 99% training accuracy and 62% on photos sent by farmers.
When they look at the mistakes, they find the model learned that "diseased" photos were mostly taken against a white sheet, and "healthy" ones against soil. It learned the background, not the disease.
The fixes: collect 800 more photos from real farms, add random crops, rotations and brightness changes, switch to a pretrained EfficientNet-B0, and use early stopping. Farmer-photo accuracy rises to 88%, and the training–validation gap shrinks from 37 points to 5.
Follow-up questions to expect
- "How do you tell overfitting from underfitting?" — Compare training and validation metrics. Big gap with good training score is overfitting; both poor is underfitting.
- "Why do huge LLMs not overfit badly?" — They are trained on so much data that they see most text only once or a few times. Overfitting depends on the ratio of capacity to data, not size alone.
- "Can the validation set itself be overfitted?" — Yes. If you tune many hyperparameters against it, you slowly fit its quirks. That is why you keep a separate test set and use it once.