Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What role do weights and bias play in a neural network? How are the weights initialized?


What you need to know

The role of weights and bias, briefly

  • A weight multiplies one input. Large magnitude means large influence; the sign says whether the input pushes the neuron up or down.
  • A bias is added after the weighted sum. It moves the neuron's threshold, so the neuron can be active even when inputs are zero, or stay quiet when they are large.

Why the starting scale matters

Each layer multiplies its input by a weight matrix. If weights are too large, the values grow at every layer and explode after 20 layers. If too small, they shrink towards zero. The same happens to gradients on the way back. Good initialisation keeps the variance of activations roughly the same from layer to layer.

The standard schemes

SchemeVariance of weightsBest withExample: Linear(512, 256)
Xavier / Glorot (normal)2 / (fan_in + fan_out)tanh, sigmoidstd ≈ 0.051
He / Kaiming (normal)2 / fan_inReLU, LeakyReLU, GELUstd ≈ 0.0625
PyTorch nn.Linear defaultuniform in ±1/sqrt(fan_in)general defaultrange ±0.044

He uses a factor of 2 because ReLU sets about half the values to zero, which halves the variance; doubling the weights' variance makes up for it.

Python
import torch.nn as nnlayer = nn.Linear(512, 256)nn.init.kaiming_normal_(layer.weight, nonlinearity="relu")   # He initnn.init.zeros_(layer.bias)# for a tanh layer instead:# nn.init.xavier_normal_(layer.weight)

PyTorch's layers already have a sensible default, but it is not exactly He initialisation, and its biases are small random values, not zeros. For custom deep networks, explicitly using He for ReLU is a common, safe choice. Pretrained models skip this question entirely, because their weights come from the checkpoint.

Other techniques to know

  • Orthogonal initialisation — often used for RNN weight matrices.
  • Zero-initialising the last layer of each residual block — each block starts as the identity, which makes very deep ResNets and transformers train more smoothly.
  • Normalisation layers (batch norm, layer norm) make networks much less sensitive to the initial scale.

A real-life example

A team builds a 20-layer MLP from scratch to classify sensor readings from farm irrigation pumps as normal or faulty. They initialise all weights from a normal distribution with standard deviation 1.

The first forward pass gives activations in the millions at layer 20 and the loss is NaN after 3 batches. Each Linear(256, 256) layer multiplied the scale by roughly sqrt(256 × 1 / 2) ≈ 11 after ReLU, so 20 layers multiplied it by about 11^20.

They switch to He initialisation (standard deviation sqrt(2/256) ≈ 0.088). Now activations at layer 20 have roughly the same scale as at layer 1, the loss falls steadily from the first epoch, and the model reaches 96% validation accuracy. Nothing changed except the starting values.

Follow-up questions to expect

  • "Why not initialise with very small values, like 0.0001?" — Activations and gradients shrink at every layer and vanish in deep networks, so early layers barely learn.
  • "Does initialisation matter with batch norm?" — Less, because batch norm rescales activations at every layer, but it still affects early training speed.
  • "How do you initialise the bias of the output layer for an imbalanced problem?" — Set it so the initial prediction equals the base rate; for 1% positives, bias = ln(0.01 / 0.99) ≈ −4.6. This avoids a large loss in the first steps.