Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the role of layers in a deep neural network?


What each layer of a speech-command model learnsInput: spectrogram, 40 bands by 100 time stepsEarly conv layers: short sounds like s, p, tMiddle layers: syllable-like patternsLast hidden layer: 64 numbers, one cluster per wordOutput: 12 logits, then softmax
Each layer's job is to pull the classes further apart, so the output layer only has to draw straight lines between clusters.

What you need to know

Three kinds of layers

  • Input layer — not really a computation; it is the raw features (pixels, audio frames, word IDs).
  • Hidden layers — every layer between input and output. This is where representation learning happens.
  • Output layer — sized and activated for the task: one neuron per class with softmax for multi-class, one neuron with sigmoid for yes/no, one linear neuron for a number.

A hierarchy of features

In a trained vision network you can look at what makes each layer's neurons fire:

  1. Layer 1 fires on edges at different angles and colour blobs.
  2. Layer 2–3 fires on corners, stripes, simple textures.
  3. Middle layers fire on patterns like spots, veins, fur.
  4. Late layers fire on object parts: a leaf's tip, a dog's ear.

Nobody designs this. It emerges because building complex features from simpler ones is the most efficient way to reduce the loss.

Why the non-linearity is essential

A linear layer computes W x + b. Two linear layers in a row give W2 (W1 x + b1) + b2 = (W2 W1) x + (W2 b1 + b2), which is again one linear layer. You can check it:

Python
import torch, torch.nn as nna, c = nn.Linear(10, 20), nn.Linear(20, 5)x = torch.randn(8, 10)W = c.weight @ a.weightB = c.weight @ a.bias + c.biasprint(torch.allclose(c(a(x)), x @ W.T + B, atol=1e-6))   # True

The two-layer network gives exactly the same output as one merged layer. Put nn.ReLU() between them and no such merge exists: now the network can bend its decision boundary.

A real-life example

A speech-commands model hears one second of audio and must output one of 12 words.

  • The input is a spectrogram: 40 frequency bands × 100 time steps.
  • Early convolution layers learn short sound events: a burst of noise ("s"), a vowel-like tone, a sharp stop ("p", "t").
  • Middle layers combine these into syllable-like patterns.
  • The final hidden layer produces a 64-number representation in which all the "yes" clips sit close together, far from the "no" clips.
  • The output layer is 12 neurons with softmax. Given a good representation, it only has to draw straight lines between clusters.

When the team plots the final hidden layer with t-SNE, the 12 words form 12 clear clusters. When they plot the raw spectrograms, everything overlaps. The layers did the work of pulling the classes apart.

Follow-up questions to expect

  • "What does 'representation learning' mean?" — The hidden layers learn useful features automatically, instead of a human designing them.
  • "Why is the last hidden layer useful on its own?" — It is an embedding. Teams reuse it for search, clustering, or as input to a small new classifier: that is the basis of transfer learning.
  • "Do all layers learn at the same speed?" — No. In deep plain networks, early layers often learn more slowly because their gradients are smaller. Normalisation and residual connections help.