Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Why do we use CNNs for image data instead of feedforward neural networks?


What you need to know

The first CNN section covered the parameter maths and weight sharing. Here is the angle interviewers push on next: inductive bias.

What the feedforward network has to learn from scratch

A dense layer treats the flattened image as 150,528 unrelated numbers. It does not know that pixel 1 and pixel 2 are neighbours, or that pixel 1 and pixel 225 are neighbours vertically. It could learn this from data, but it has to learn it separately for every position. That takes an enormous amount of data, and with millions of weights and limited images it memorises instead.

A CNN gets these facts for free from its structure:

  • Locality: each filter only sees a small patch.
  • Weight sharing: one filter is reused at every position.
  • Hierarchy: stacking layers grows the area each neuron sees, from 3×3 pixels to the whole image.

The data-efficiency argument

Fewer free parameters, each constrained by a sensible assumption, means the model can reach good accuracy from thousands of images rather than millions. That is the real reason CNNs won on images, more than raw speed.

When the answer changes

  • Tiny, centred images such as 28×28 digits: an MLP can reach good accuracy, though a small CNN still does better.
  • Vision transformers (ViTs) have weaker built-in assumptions than CNNs. They can match or beat CNNs, but typically need large-scale pretraining or heavy augmentation to make up for it.
  • Tabular data: no spatial structure, so CNNs bring nothing; use gradient-boosted trees or an MLP.

A real-life example

A diagnostic-imaging start-up has 12,000 labelled chest X-rays at 1024×1024 grayscale. That is 1,048,576 inputs per image. A single dense layer of 512 neurons would need over 536 million weights, about 45,000 weights per training image. Such a model memorises the training set, including hospital-specific markers such as a text label in the corner, and fails on X-rays from a new hospital.

A ResNet-style CNN, pretrained on natural images and fine-tuned, has about 25 million parameters in total, almost all in shared filters. It learns what lung opacities look like regardless of where they sit. The team still resizes images to 512×512 to fit GPU memory, and checks with heatmaps that the model looks at the lungs and not the corner labels.

Follow-up questions to expect

  • "Is a CNN translation-invariant?" — Convolution itself is translation-equivariant: shift the input and the feature map shifts. Pooling and global pooling add partial invariance. Large shifts, rotations and scale changes still need augmentation.
  • "What is a receptive field?" — The region of the input that affects one neuron. Two stacked 3×3 convolutions give a 5×5 receptive field, three give 7×7.
  • "Would a CNN help on shuffled pixels?" — No. If you permute the pixels the same way for every image, the local structure is gone and the CNN's advantage disappears; a dense network would do about as well.