Course Content
Computer Vision Fundamentals
3 sections · 8 lessons
Architectures (LeNet, AlexNet, VGG, ResNet)
In 2014 a reasonable person could look at the progress in image recognition and predict the future confidently: networks were getting deeper, and deeper networks were getting better. So build a 56-layer network and it should beat the 20-layer one.
Researchers tried exactly that. The 56-layer network was worse. Not worse on the test set — worse on the training set, which rules out overfitting entirely. A model that cannot even fit the data it has memorisation capacity for is failing at optimisation, not generalisation.
This was genuinely strange, because the 56-layer network provably can represent everything the 20-layer network represents. Take the 20-layer solution and make the extra 36 layers compute the identity function — pass input straight through — and you have an exact copy. So a solution at least as good is definitely in there. Gradient descent just could not find it.
The history of convolutional architectures is largely the history of people running into walls like that one and finding a way round. Reading them in order is not nostalgia: each design decision you now take for granted was somebody's answer to a specific failure, and knowing the failure tells you when the answer stops applying.
LeNet-5, 1998: the template
Yann LeCun's network read handwritten digits on bank cheques; its predecessors read ZIP codes for the US postal service. It ran on 1998 hardware — no GPUs, roughly a millionth of the compute a modern accelerator provides — and it worked in production.
Input 32x32x1 (grayscale digit) Conv 5x5, 6 filters -> 28x28x6 AvgPool 2x2 -> 14x14x6 Conv 5x5, 16 filters -> 10x10x16 AvgPool 2x2 -> 5x5x16 Flatten -> 400 Dense 120 -> Dense 84 -> Dense 10About 60,000 parameters — small enough to fit comfortably in a modern phone's cache.
What matters is that LeNet established the template every subsequent architecture follows: alternating convolution and downsampling to build up features, then dense layers to classify. Spatial size shrinks (32 → 28 → 14 → 10 → 5) while channel count grows (1 → 6 → 16). Detail is traded for abstraction.
It used tanh activations and average pooling, both of which were later replaced. And then the field went quiet for over a decade — not because the idea was wrong, but because the three things it needed did not exist yet: enough labelled data, enough compute, and a few training tricks.
AlexNet, 2012: the one that changed everything
The ImageNet competition asked models to classify 1.2 million photographs into 1000 categories. In 2011 the winning error rate was 25.8%, achieved with hand-engineered features. In 2012 AlexNet scored 15.3%. The runner-up, still using hand-engineered features, scored 26.2%.
Input 224x224x3 Conv 11x11, 96 filters, stride 4 -> 55x55x96 MaxPool 3x3, stride 2 -> 27x27x96 Conv 5x5, 256 filters, pad 2 -> 27x27x256 MaxPool 3x3, stride 2 -> 13x13x256 Conv 3x3, 384 -> Conv 3x3, 384 -> Conv 3x3, 256 MaxPool 3x3, stride 2 -> 6x6x256 Dense 4096 -> Dense 4096 -> Dense 1000Eight learned layers, roughly 60 million parameters. Structurally it is LeNet made bigger. The wins came from four changes, each addressing a specific obstacle:
| Change | Problem it solved |
|---|---|
| ReLU instead of tanh | tanh saturates, so gradients shrink towards zero in deep stacks and training crawls. ReLU's gradient is 1 for positive inputs. Reported as roughly 6× faster to converge. |
| Training on two GPUs | The model did not fit in 3 GB. Channels were split across two cards, which is also where grouped convolution came from. |
| Dropout in the dense layers | 58 million of the 60 million parameters are in the dense layers. Without dropout they memorised the training set. |
| Heavy augmentation | Random crops and flips multiplied the effective dataset and cut overfitting further. |
The parameter distribution is the lesson worth carrying away. Count them: the final 6×6×256 feature map flattens to 9216 values, and connecting that to a 4096-unit dense layer costs 9216 × 4096 ≈ 37.7 million weights. One layer, over 60% of the model. The convolutional layers — the part doing the actual visual work — hold under 4% of the parameters.
Almost all the parameters sit in the dense head, and almost all the computation sits in the convolutional body. Optimising the wrong one wastes your effort.
VGG, 2014: uniformity as a design principle
AlexNet's kernel sizes were 11, then 5, then 3, chosen by intuition. VGG asked what happens if you use 3×3 everywhere and simply go deeper.
VGG-16, in blocks (all convs 3x3, pad 1; all pools 2x2, stride 2) [Conv 64, Conv 64 ] Pool 224 -> 112 [Conv 128, Conv 128] Pool 112 -> 56 [Conv 256 x3 ] Pool 56 -> 28 [Conv 512 x3 ] Pool 28 -> 14 [Conv 512 x3 ] Pool 14 -> 7 Flatten 7x7x512 = 25088 Dense 4096 -> Dense 4096 -> Dense 1000The justification for 3×3 everywhere is the stacking argument. Two 3×3 convolutions see the same 5×5 region as one 5×5 convolution, and three see the same 7×7 region as one 7×7. But:
| Option | Receptive field | Weights per channel pair | ReLUs |
|---|---|---|---|
| One 7×7 conv | 7×7 | 49 | 1 |
| Three 3×3 convs | 7×7 | 27 | 3 |
Fewer parameters and more non-linearity, for the same field of view. That argument won, and 3×3 has been the default ever since.
VGG-16 reached 7.3% ImageNet error with 138 million parameters. And here the flaw becomes glaring: the flatten step produces 25,088 values, and the first dense layer needs 25,088 × 4096 ≈ 102.8 million weights. Three-quarters of the entire model is one layer that does no spatial reasoning at all.
VGG remains useful precisely because it is so regular — it is easy to read, easy to modify, and its intermediate features are still used for perceptual loss functions in image generation. But nobody trains it for deployment. It is slow and enormous for its accuracy.
ResNet, 2015: solving the degradation problem
Back to the puzzle at the top. Deeper networks were performing worse on training data, which is an optimisation failure, not overfitting. This was called the degradation problem.
The diagnosis: layers find it surprisingly hard to learn the identity function. To pass input through unchanged, a layer's weights must land in a very specific configuration, and gradient descent has no particular reason to find it. Meanwhile the gradient signal reaching early layers of a very deep stack has been multiplied through dozens of Jacobians, shrinking it towards nothing.
The fix
Kaiming He and colleagues changed what each block is asked to learn. Instead of a block computing an output H(x) directly, it computes a residual F(x), and the block's output is:
The input is added straight to the output through a skip connection that bypasses the layers entirely.
x ─────────────────────┐ │ │ (identity shortcut) Conv 3x3 → BN → ReLU │ │ │ Conv 3x3 → BN │ │ │ + ◄────────────────────┘ │ ReLU │ yNow consider what it costs to represent the identity. The block must produce y=x, so F(x) must equal zero — and driving weights towards zero is precisely what weight decay and gradient descent do naturally. The easy default became "change nothing", and any layer that has nothing useful to add can effectively switch itself off. Adding layers can no longer hurt.
The gradient story is equally important. Differentiating y=F(x)+x with respect to x gives:
That +1 is the whole point. Even if ∂F/∂x collapses to nearly zero, the gradient still flows back through the shortcut undiminished. Skip connections give gradients a clear path from the loss all the way to the first layer, however deep the network.
The skip connection makes "do nothing" the cheapest option and gives gradients an unobstructed route home. Those two properties are what made networks of 100+ layers trainable.
The result: a single ResNet-152 reached about 4.5% top-5 error, an ensemble of ResNets won ILSVRC 2015 with 3.57% — below the roughly 5% human estimate for that task — and a 1202-layer version trained on CIFAR-10 without diverging.
The bottleneck block
Deeper ResNets replace the two 3×3 layers with a three-layer sandwich, and the arithmetic is worth working through.
For 256 channels, two stacked 3×3 convolutions cost:
2×(256×3×3×256)=1,179,648 parametersThe bottleneck instead does 1×1 down to 64 channels, 3×3 at 64, then 1×1 back up to 256:
Layer Shape change Parameters 1×1 conv 256 → 64 16,384 3×3 conv 64 → 64 36,864 1×1 conv 64 → 256 16,384 Total 69,632 Seventeen times cheaper, while keeping the expensive 3×3 spatial operation. The 1×1 layers cost almost nothing because they have no spatial extent — they are just learned mixing across channels. This pattern, squeeze the channels, do the spatial work cheaply, expand again, appears everywhere in efficient architecture design.
When the shortcut cannot be an identity
A skip connection adds x to F(x), which requires matching shapes. When a block changes channel count or downsamples, they do not match. Two options:
- Projection shortcut — a 1×1 convolution with the appropriate stride on the shortcut path, to reshape x. This is the standard choice and adds a modest number of parameters.
- Zero-padding shortcut — downsample by striding and pad the extra channels with zeros. Free, but slightly worse.
1import torch.nn as nn23class BasicBlock(nn.Module):4 def __init__(self, c_in, c_out, stride=1):5 super().__init__()6 self.conv1 = nn.Conv2d(c_in, c_out, 3, stride=stride, padding=1, bias=False)7 self.bn1 = nn.BatchNorm2d(c_out)8 self.conv2 = nn.Conv2d(c_out, c_out, 3, stride=1, padding=1, bias=False)9 self.bn2 = nn.BatchNorm2d(c_out)10 self.relu = nn.ReLU(inplace=True)1112 # Projection only when the shapes would not line up13 self.shortcut = nn.Identity()14 if stride != 1 or c_in != c_out:15 self.shortcut = nn.Sequential(16 nn.Conv2d(c_in, c_out, 1, stride=stride, bias=False),17 nn.BatchNorm2d(c_out),18 )1920 def forward(self, x):21 out = self.relu(self.bn1(self.conv1(x)))22 out = self.bn2(self.conv2(out))23 out = out + self.shortcut(x) # add BEFORE the final activation24 return self.relu(out)Note the ordering: the addition happens before the final ReLU, not after. Putting ReLU on the shortcut path would clip negative values out of the identity signal and undo much of the benefit.
The other structural win
ResNet replaced the flatten-plus-dense head with global average pooling: average each of the 512 (or 2048) final feature maps down to a single number, then one linear layer to the classes. ResNet-50's entire classifier head is 2048 × 1000 ≈ 2 million parameters, against VGG's 102.8 million. That single change is why ResNet-50 has fewer parameters than VGG-16 while being far deeper and considerably more accurate.
Side by side
| Model | Year | Layers | Parameters | ImageNet top-5 error | Key contribution |
|---|---|---|---|---|---|
| LeNet-5 | 1998 | 7 | 60 K | — (MNIST) | The conv-pool-dense template |
| AlexNet | 2012 | 8 | 60 M | 15.3% | ReLU, dropout, GPU training at scale |
| VGG-16 | 2014 | 16 | 138 M | 7.3% | Uniform 3×3 stacks; depth over width |
| ResNet-50 | 2015 | 50 | 25 M | 5.3% | Skip connections; bottleneck blocks |
| ResNet-152 | 2015 | 152 | 60 M | 4.5% (3.6% ensemble) | Depth that finally pays off |
Read the parameter column against the error column. ResNet-50 has 18% of VGG-16's parameters and lower error. Parameter count is not a proxy for capability; architecture is. The gains came from better structure, not more weights.
What this means when you pick a backbone
For almost any classification task, a pretrained ResNet-50 is the sensible first baseline, for a specific reason: it is accurate enough that architecture will not be your bottleneck, small enough to fine-tune on one GPU, and so widely used that pretrained weights, tutorials and debugging advice exist for every framework and every odd situation you will hit. Newer backbones — ConvNeXt, EfficientNetV2, or vision transformers with self-supervised weights such as DINOv2 — often score a few points higher, so try one once the baseline works; but reach for anything more exotic only when you have measured a concrete problem it solves.
When you do need to move, move for a measured reason. Too slow on a phone means you want depthwise separable convolutions, which trade a little accuracy for roughly an order of magnitude fewer parameters. Too little data means a smaller backbone such as ResNet-18, since a large model on a few thousand images will overfit no matter how well it is regularised. Objects too small to detect usually means the receptive field or the feature resolution is wrong, not that the backbone is wrong.
Three ideas from this history transfer to any network you design, whatever the domain. Prefer small kernels stacked deep over large kernels used shallow, because you get more non-linearity for fewer weights. Add skip connections to anything deeper than about ten layers, because they cost almost nothing and make optimisation dramatically easier. And check where your parameters actually live before optimising: if three-quarters of your model is one dense layer, replacing it with global average pooling is a far bigger win than any amount of tuning elsewhere.