Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are deep and shallow networks? Which one is better and why?


What you need to know

Shallow vs deep

  • A shallow network has an input layer, one hidden layer (sometimes two), and an output layer.
  • A deep network has many hidden layers: ResNet-50 has 50 weight layers, and large language models have dozens of transformer blocks.

Why depth helps: composition

Real-world data is built in levels. Pixels form edges, edges form shapes, shapes form a leaf with spots. A deep network matches this structure: each layer builds on the one below, and a feature learned once (say, "a curved edge") is reused by many higher features.

A shallow network has to jump from raw pixels to the answer in one step. For some functions, researchers have proved that a shallow network needs exponentially more neurons than a deep one to match it. In practice, deep networks reach the same accuracy with fewer parameters.

What depth costs

  • Training difficulty — gradients can shrink (vanish) or blow up (explode) as they pass back through many layers. Residual connections, normalisation and ReLU-family activations fix most of this.
  • More data and compute — more layers usually means more parameters to fit.
  • Latency — layers run one after another, so depth adds inference time.

Shallow network

  • One or two hidden layers
  • Learns directly from input to output
  • Fast to train, needs less data
  • Good for small tabular problems

Deep network

  • Many hidden layers
  • Builds features on top of features
  • Needs more data, compute and care
  • Best for images, audio, text

A real-life example

A team classifies 12 spoken commands ("yes", "no", "up", "down", ...) from one-second audio clips.

  • Model A: a shallow MLP with one hidden layer of 4,096 neurons on the flattened spectrogram. About 17 million parameters. Test accuracy around 80%, and it overfits quickly.
  • Model B: a small CNN with 6 convolutional layers. About 250,000 parameters. Test accuracy around 95%.

Model B wins with 70 times fewer parameters because its early layers learn short sound patterns, and later layers combine them into syllables and words. Model A has to learn every word at every time position separately.

The same team's customer-churn model on 15 tabular columns works best as a shallow network or, better, gradient boosting. There is no hierarchy in those columns for depth to exploit.

Follow-up questions to expect

  • "If one hidden layer is enough in theory, why go deep?" — Because "enough" may mean an impossibly wide layer, and training may not find the solution. Depth gives the same power with fewer parameters and better generalisation.
  • "Can a network be too deep?" — Yes. Plain networks past about 20–30 layers started training worse, not just slower. Residual connections fixed that and made 100+ layer networks trainable.
  • "Is wider or deeper better?" — It depends on the data. Width adds features at one level; depth adds levels of abstraction. Modern architectures scale both together.