Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is model capacity?


Three fraud models, 3,000 fraud cases410.780.76underfit6900.860.84good fit284,6740.990.71overfitParamsTrain recallVal recallVerdictLogistic regressionSmall net + decayLarge net
Read the gap between training and validation, not the validation number alone: the gap tells you which way to move capacity.

What you need to know

What sets capacity

  • Number of parameters — more weights, more ways to bend the function.
  • Depth and width — more layers and more neurons per layer.
  • Activation functions — without non-linear activations, even a deep network is only a linear model.
  • Training time — a network trained for longer uses more of its capacity. This is why early stopping works as a form of capacity control.

Capacity and the bias–variance trade-off

Too little capacityRight capacityToo much, uncontrolled
Training errorHighLowNear zero
Validation errorHighLowHigh, rising
NameUnderfitting (high bias)Good fitOverfitting (high variance)

Bias is error from wrong assumptions: the model is too simple to follow the pattern. Variance is error from being too sensitive to the specific training set: change the data slightly and the model changes a lot.

How practice works today

Modern deep learning often uses more capacity than needed and then controls it with weight decay, dropout, data augmentation, early stopping and lots of data. Large, well-regularised networks often generalise better than small ones. So "reduce capacity" is not the only fix for overfitting; "add regularisation or data" often works better.

You can count parameters to get a rough feel for capacity:

Python
import torch.nn as nnsmall = nn.Sequential(nn.Linear(40, 16), nn.ReLU(), nn.Linear(16, 2))large = nn.Sequential(nn.Linear(40, 512), nn.ReLU(),                      nn.Linear(512, 512), nn.ReLU(), nn.Linear(512, 2))count = lambda m: sum(p.numel() for p in m.parameters())print(count(small), count(large))   # 690  284674

Each Linear(a, b) has a*b weights plus b biases. The large model has about 400 times more parameters, so it can fit far more complex functions, and far more noise.

A real-life example

A bank has only 3,000 labelled credit-card fraud cases and 40 features per transaction.

  • A logistic regression (41 parameters) reaches 0.78 recall on training and 0.76 on validation. Both similar, both mediocre: underfitting.
  • The 284,674-parameter network above reaches 0.99 training recall and 0.71 validation recall. It has memorised the 3,000 fraud rows: overfitting.
  • A 690-parameter network with weight decay and early stopping reaches 0.86 training and 0.84 validation. Enough capacity, controlled.

The engineer reads the gap between training and validation, not just the validation number, to decide which way to move.

Follow-up questions to expect

  • "How do you measure capacity?" — Parameter count is a rough guide. Formal measures like VC dimension exist, but in practice I compare training and validation curves.
  • "How do you increase capacity if you are underfitting?" — Add layers or neurons, train longer, reduce regularisation, or use a better architecture for the data type.
  • "Can a huge model generalise well?" — Yes, with enough data and regularisation. Large pretrained models fine-tuned on small datasets often generalise very well.