Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is model capacity?
What you need to know
What sets capacity
- Number of parameters — more weights, more ways to bend the function.
- Depth and width — more layers and more neurons per layer.
- Activation functions — without non-linear activations, even a deep network is only a linear model.
- Training time — a network trained for longer uses more of its capacity. This is why early stopping works as a form of capacity control.
Capacity and the bias–variance trade-off
| Too little capacity | Right capacity | Too much, uncontrolled | |
|---|---|---|---|
| Training error | High | Low | Near zero |
| Validation error | High | Low | High, rising |
| Name | Underfitting (high bias) | Good fit | Overfitting (high variance) |
Bias is error from wrong assumptions: the model is too simple to follow the pattern. Variance is error from being too sensitive to the specific training set: change the data slightly and the model changes a lot.
How practice works today
Modern deep learning often uses more capacity than needed and then controls it with weight decay, dropout, data augmentation, early stopping and lots of data. Large, well-regularised networks often generalise better than small ones. So "reduce capacity" is not the only fix for overfitting; "add regularisation or data" often works better.
You can count parameters to get a rough feel for capacity:
1import torch.nn as nn23small = nn.Sequential(nn.Linear(40, 16), nn.ReLU(), nn.Linear(16, 2))4large = nn.Sequential(nn.Linear(40, 512), nn.ReLU(),5 nn.Linear(512, 512), nn.ReLU(), nn.Linear(512, 2))67count = lambda m: sum(p.numel() for p in m.parameters())8print(count(small), count(large)) # 690 284674Each Linear(a, b) has a*b weights plus b biases. The large model has about 400 times more parameters, so it can fit far more complex functions, and far more noise.
A real-life example
A bank has only 3,000 labelled credit-card fraud cases and 40 features per transaction.
- A logistic regression (41 parameters) reaches 0.78 recall on training and 0.76 on validation. Both similar, both mediocre: underfitting.
- The 284,674-parameter network above reaches 0.99 training recall and 0.71 validation recall. It has memorised the 3,000 fraud rows: overfitting.
- A 690-parameter network with weight decay and early stopping reaches 0.86 training and 0.84 validation. Enough capacity, controlled.
The engineer reads the gap between training and validation, not just the validation number, to decide which way to move.
Follow-up questions to expect
- "How do you measure capacity?" — Parameter count is a rough guide. Formal measures like VC dimension exist, but in practice I compare training and validation curves.
- "How do you increase capacity if you are underfitting?" — Add layers or neurons, train longer, reduce regularisation, or use a better architecture for the data type.
- "Can a huge model generalise well?" — Yes, with enough data and regularisation. Large pretrained models fine-tuned on small datasets often generalise very well.