Course Content
Deep Learning Essentials
13 sections · 61 lessons
Can a neural network be trained by initializing all the weights to 0 or a constant?
What you need to know
Why identical neurons stay identical
Take a hidden layer where every weight is 0.5. For any input, every hidden neuron computes the same weighted sum, so the same activation. The output layer's weights are also all equal, so each hidden neuron has the same effect on the loss. Backpropagation gives each one the same gradient, the optimiser applies the same update, and after the update they are still identical. This repeats every step.
Seeing it happen
1import torch, torch.nn as nn23X, y = torch.randn(64, 4), torch.randint(0, 2, (64,)).float()45def train(init):6 net = nn.Sequential(nn.Linear(4, 3), nn.ReLU(), nn.Linear(3, 1))7 for m in (net[0], net[2]):8 init(m.weight); nn.init.zeros_(m.bias)9 opt = torch.optim.SGD(net.parameters(), lr=0.1)10 for _ in range(100):11 loss = nn.functional.binary_cross_entropy_with_logits(net(X).squeeze(-1), y)12 opt.zero_grad(); loss.backward(); opt.step()13 return net[0].weight.data1415print(train(nn.init.zeros_)) # still all zeros16print(train(lambda w: nn.init.constant_(w, 0.5))) # 3 identical rows17print(train(lambda w: nn.init.kaiming_normal_(w, nonlinearity="relu"))) # 3 different rowsAfter 100 steps, one run gave:
- All zeros: the first layer is still exactly zero. Hidden outputs are
ReLU(0) = 0, and the output weights are zero, so no gradient reaches the hidden layer. Only the output bias learns. - Constant 0.5: the weights changed, but all three rows are still exactly the same four numbers. Three neurons, one feature.
- He: three different rows. Three neurons learned three different features.
The clarifications that show depth
- Biases can start at zero. Symmetry is already broken by the random weights.
- Single-layer models are fine at zero. Logistic or linear regression has no hidden neurons to be symmetric, so it trains normally from zero.
- Some zeros are deliberate. Setting the last layer of each residual block to zero is safe, because the other layers in the block are random, so symmetry is still broken.
A real-life example
A developer writes a custom fraud model from scratch in NumPy for a coding test and initialises weights with np.zeros. The loss drops a little in the first few steps and then goes flat. Checking the model, every prediction equals the fraud base rate: only the output bias learned.
They inspect the first layer: 64 neurons, all weights exactly 0. They switch to np.random.randn(n_in, n_out) * np.sqrt(2 / n_in), He initialisation. The loss keeps falling and the hidden neurons learn 64 different patterns, from "high amount" to "new device at night".
In the interview debrief, explaining the fix with the word "symmetry" turned a bug into a strong answer.
Follow-up questions to expect
- "What about initialising with the same random value for every weight?" — Same problem. What matters is that weights differ from each other, not that they are random-looking.
- "Would all-zero weights work with sigmoid?" — Still broken. Hidden outputs are all 0.5, so after the output weights move, every hidden neuron gets the same gradient and they stay identical.
- "Why is zero fine for biases but not weights?" — Bias gradients depend on the neuron's error signal, which already differs between neurons once the weights are random.