Course Content
Deep Learning Essentials
13 sections · 61 lessons
Why are Convolutional Neural Networks better suited for image data than fully connected networks?
What you need to know
The problem with fully connected layers on images
A fully connected (dense) layer connects every input to every neuron. A 224×224 RGB image has 224 × 224 × 3 = 150,528 input values. One dense layer with 1,000 neurons needs:
150,528 inputs × 1,000 neurons + 1,000 biases = 150,529,000 parametersThat is one layer. It also throws away the image's layout: after flattening, pixel (10, 10) and pixel (10, 11) are just two unrelated numbers in a long list. The network must learn separately that a cat's ear at the top-left and at the bottom-right are the same thing.
The three ideas that make CNNs work
- Local connectivity. Each neuron in a conv layer looks at a small patch, such as 3×3 pixels. Edges, corners and textures are local, so this is enough for early layers. Deeper layers see larger areas because they combine patches from the layer below; this area is called the receptive field.
- Weight sharing. The same filter is applied at every position. A filter that finds vertical edges is learned once and used everywhere.
- Translation equivariance, then some invariance. Because the filter is shared, if an object moves 10 pixels right, its feature map moves 10 pixels right too. That is equivariance. Pooling and the final global pooling layer then add partial invariance: the prediction stays the same even if the object is somewhat shifted.
Parameter count, side by side
1import torch.nn as nn23fc = nn.Linear(224 * 224 * 3, 1000)4conv = nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3, padding=1)56count = lambda m: sum(p.numel() for p in m.parameters())7print(f"{count(fc):,} {count(conv):,}") # 150,529,000 1,792The conv layer has 64 filters, each 3×3×3 = 27 weights plus 1 bias, so 64 × 28 = 1,792 parameters. It still produces a 64×224×224 output, a full map for every filter. Fewer parameters means less to learn, so the model needs far less data to generalise.
Fully connected on pixels
- Every pixel connects to every neuron
- About 150 million weights in one layer
- Loses 2D structure after flattening
- Must relearn each pattern at each position
Convolutional layer
- Each neuron sees a small local patch
- 1,792 weights for 64 filters
- Keeps the height and width layout
- One filter detects its pattern anywhere
A real-life example
A startup diagnoses rice-leaf diseases (blast, brown spot, bacterial blight) from 5,000 phone photos resized to 224×224.
Their first try flattens the image into an MLP with two hidden layers. With 1,000 neurons in the first hidden layer, it has over 150 million parameters. It reaches 98% training accuracy but 58% on validation, and when a farmer holds the leaf slightly off-centre, predictions change. It learned "brown pixels at position (120, 87)", not "brown spot lesion".
A small CNN with five conv blocks has under 2 million parameters and reaches 90% validation accuracy. When the lesions appear at a different place on the leaf, the same filters still fire, because they are shared across all positions. Fine-tuning a pretrained ResNet pushes it to 95%.
Follow-up questions to expect
- "Are CNNs fully translation invariant?" — No. Convolution is equivariant; pooling and global average pooling give only partial invariance. Large shifts, rotations and scale changes still need data augmentation.
- "Where do the parameters of a CNN live?" — In conv filters, their biases, batch-norm layers and any fully connected layers at the end. Pooling layers have none.
- "Are CNNs still used now that vision transformers exist?" — Yes. On small and medium datasets and on phones, CNNs are often more data-efficient and faster. Vision transformers tend to win with very large datasets or pretraining.