Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the role of layers in a deep neural network?
What you need to know
Three kinds of layers
- Input layer — not really a computation; it is the raw features (pixels, audio frames, word IDs).
- Hidden layers — every layer between input and output. This is where representation learning happens.
- Output layer — sized and activated for the task: one neuron per class with softmax for multi-class, one neuron with sigmoid for yes/no, one linear neuron for a number.
A hierarchy of features
In a trained vision network you can look at what makes each layer's neurons fire:
- Layer 1 fires on edges at different angles and colour blobs.
- Layer 2–3 fires on corners, stripes, simple textures.
- Middle layers fire on patterns like spots, veins, fur.
- Late layers fire on object parts: a leaf's tip, a dog's ear.
Nobody designs this. It emerges because building complex features from simpler ones is the most efficient way to reduce the loss.
Why the non-linearity is essential
A linear layer computes W x + b. Two linear layers in a row give W2 (W1 x + b1) + b2 = (W2 W1) x + (W2 b1 + b2), which is again one linear layer. You can check it:
1import torch, torch.nn as nn23a, c = nn.Linear(10, 20), nn.Linear(20, 5)4x = torch.randn(8, 10)5W = c.weight @ a.weight6B = c.weight @ a.bias + c.bias7print(torch.allclose(c(a(x)), x @ W.T + B, atol=1e-6)) # TrueThe two-layer network gives exactly the same output as one merged layer. Put nn.ReLU() between them and no such merge exists: now the network can bend its decision boundary.
A real-life example
A speech-commands model hears one second of audio and must output one of 12 words.
- The input is a spectrogram: 40 frequency bands × 100 time steps.
- Early convolution layers learn short sound events: a burst of noise ("s"), a vowel-like tone, a sharp stop ("p", "t").
- Middle layers combine these into syllable-like patterns.
- The final hidden layer produces a 64-number representation in which all the "yes" clips sit close together, far from the "no" clips.
- The output layer is 12 neurons with softmax. Given a good representation, it only has to draw straight lines between clusters.
When the team plots the final hidden layer with t-SNE, the 12 words form 12 clear clusters. When they plot the raw spectrograms, everything overlaps. The layers did the work of pulling the classes apart.
Follow-up questions to expect
- "What does 'representation learning' mean?" — The hidden layers learn useful features automatically, instead of a human designing them.
- "Why is the last hidden layer useful on its own?" — It is an embedding. Teams reuse it for search, clustering, or as input to a small new classifier: that is the basis of transfer learning.
- "Do all layers learn at the same speed?" — No. In deep plain networks, early layers often learn more slowly because their gradients are smaller. Normalisation and residual connections help.