Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What problem does convolution solve in computer vision tasks?


Feature map from a Sobel vertical-edge filter036360036360036360cols 0-2cols 1-3cols 2-4cols 3-5Input: 5 by 6 pixels, value 0 on the left three columns and 9 on the right three.
The feature map is a map of where the pattern is: 36 wherever the 3 by 3 window straddles the dark-to-bright edge, 0 on flat regions.

What you need to know

What a convolution computes

Take a small grayscale image and a 3×3 filter. Put the filter on the top-left 3×3 patch, multiply each pixel by the filter weight on top of it, add the nine products (plus a bias), and write the result into the output. Move one pixel right and repeat. When you reach the end of the row, move down.

Here is a real vertical-edge filter (Sobel) applied to an image that is dark on the left and bright on the right:

Python
import torchimport torch.nn.functional as Fimg = torch.tensor([[0, 0, 0, 9, 9, 9]] * 5, dtype=torch.float32).view(1, 1, 5, 6)sobel_x = torch.tensor([[-1, 0, 1],                        [-2, 0, 2],                        [-1, 0, 1]], dtype=torch.float32).view(1, 1, 3, 3)print(F.conv2d(img, sobel_x))# tensor([[[[ 0., 36., 36.,  0.],#           [ 0., 36., 36.,  0.],#           [ 0., 36., 36.,  0.]]]])

The output is large (36) only where the filter sits across the dark-to-bright boundary, and 0 over flat regions. That output grid is a feature map: a map of "where is there a vertical edge". The 5×6 input gave a 3×4 output because a 3×3 filter fits in 3 × 4 positions without padding.

(Strictly, deep learning libraries compute cross-correlation: they do not flip the filter as the mathematical convolution does. Since the filter weights are learned, the difference does not matter.)

The problems it solves

  • Too many parameters — a 3×3 filter is 9 weights per input channel, reused at every position, instead of one weight per pixel per neuron.
  • Loss of spatial structure — the output is still a 2D grid, so the next layer knows which features are neighbours.
  • Position dependence — the same filter finds the pattern anywhere. If the edge moves one pixel left, the 36s in the feature map move one position left.
  • Hand-crafted features — Sobel, Gabor and HOG filters were designed by people. A CNN learns hundreds of filters, tuned to its task. In a trained network, many first-layer filters look like edge and colour detectors, similar to the classic ones.

Stacking builds a hierarchy

One conv layer finds simple local patterns. The next layer convolves over those feature maps, so it finds patterns of patterns: two edges meeting make a corner, many small blobs make a texture. After several layers, each neuron responds to a large area and complex shapes. This is how a CNN goes from pixels to "this leaf has late blight".

A real-life example

A textile mill inspects fabric rolls for defects such as holes, stains and broken threads. Their old system used hand-tuned Sobel and threshold rules. It worked on plain white cotton but produced false alarms on striped or printed fabric, because the stripes are also edges.

They train a CNN on 20,000 labelled image patches. The first-layer filters learn edges, as expected. But middle layers learn something the rules never could: the regular repeating texture of each fabric. A broken thread is a break in that repetition, so it stands out even on striped cloth. False alarms fall from about 1 in 20 patches to 1 in 400, and the model runs at 30 frames per second on a small GPU next to the line.

The key difference: the old system's filters were chosen by an engineer; the CNN's filters were chosen by the loss function.

Follow-up questions to expect

  • "What does a filter with multiple input channels do?" — On RGB input, a 3×3 filter is really 3×3×3. It multiplies all three channels and sums them into one number, so each filter produces one feature map.
  • "What are stride and padding?" — Stride is how far the filter moves each step; stride 2 halves the output size. Padding adds zeros around the border so the output can keep the input's size.
  • "How does a CNN see a large object with 3×3 filters?" — Through depth and downsampling. Each layer increases the receptive field, so deep neurons see most of the image.