Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are deep and shallow networks? Which one is better and why?
What you need to know
Shallow vs deep
- A shallow network has an input layer, one hidden layer (sometimes two), and an output layer.
- A deep network has many hidden layers: ResNet-50 has 50 weight layers, and large language models have dozens of transformer blocks.
Why depth helps: composition
Real-world data is built in levels. Pixels form edges, edges form shapes, shapes form a leaf with spots. A deep network matches this structure: each layer builds on the one below, and a feature learned once (say, "a curved edge") is reused by many higher features.
A shallow network has to jump from raw pixels to the answer in one step. For some functions, researchers have proved that a shallow network needs exponentially more neurons than a deep one to match it. In practice, deep networks reach the same accuracy with fewer parameters.
What depth costs
- Training difficulty — gradients can shrink (vanish) or blow up (explode) as they pass back through many layers. Residual connections, normalisation and ReLU-family activations fix most of this.
- More data and compute — more layers usually means more parameters to fit.
- Latency — layers run one after another, so depth adds inference time.
Shallow network
- One or two hidden layers
- Learns directly from input to output
- Fast to train, needs less data
- Good for small tabular problems
Deep network
- Many hidden layers
- Builds features on top of features
- Needs more data, compute and care
- Best for images, audio, text
A real-life example
A team classifies 12 spoken commands ("yes", "no", "up", "down", ...) from one-second audio clips.
- Model A: a shallow MLP with one hidden layer of 4,096 neurons on the flattened spectrogram. About 17 million parameters. Test accuracy around 80%, and it overfits quickly.
- Model B: a small CNN with 6 convolutional layers. About 250,000 parameters. Test accuracy around 95%.
Model B wins with 70 times fewer parameters because its early layers learn short sound patterns, and later layers combine them into syllables and words. Model A has to learn every word at every time position separately.
The same team's customer-churn model on 15 tabular columns works best as a shallow network or, better, gradient boosting. There is no hierarchy in those columns for depth to exploit.
Follow-up questions to expect
- "If one hidden layer is enough in theory, why go deep?" — Because "enough" may mean an impossibly wide layer, and training may not find the solution. Depth gives the same power with fewer parameters and better generalisation.
- "Can a network be too deep?" — Yes. Plain networks past about 20–30 layers started training worse, not just slower. Residual connections fixed that and made 100+ layer networks trainable.
- "Is wider or deeper better?" — It depends on the data. Width adds features at one level; depth adds levels of abstraction. Modern architectures scale both together.