Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are weights and biases, and why are they critical in neural networks?


What you need to know

Weights

A weight is a number on the connection from one input to one neuron. A large positive weight means "this input pushes the neuron up strongly". A negative weight means it pushes down. A weight near zero means the input barely matters to this neuron.

In a layer, weights form a matrix. A PyTorch layer from 3 inputs to 4 neurons stores a 4×3 weight matrix and 4 biases:

Python
import torch.nn as nnlayer = nn.Linear(in_features=3, out_features=4)print(layer.weight.shape, layer.bias.shape)   # torch.Size([4, 3]) torch.Size([4])

Each row of the weight matrix belongs to one neuron. The layer computes x @ W.T + b for a whole batch at once.

Bias

The bias lets a neuron be active even when all inputs are zero, or stay quiet even when inputs are large. Think of the 2D boundary w1*x1 + w2*x2 + b = 0:

  • The weights w1, w2 set the line's slope and direction.
  • The bias b moves the line away from the origin.

With b = 0 the line must pass through (0, 0). If the two classes are split by a line at x1 = 5, a neuron with no bias cannot draw it.

Why they are the whole model

Training repeats one loop: compute a prediction, measure the loss, compute the gradient of the loss with respect to every weight and bias, and nudge each one to reduce the loss. When training ends, those numbers are the model. ResNet-50 has about 25 million of them. A 7-billion-parameter LLM has 7 billion.

This is also why model files are large: 7 billion parameters in 16-bit floats is about 14 GB.

A real-life example

A crop-disease classifier has, in its last layer, one neuron per class: healthy, early blight, late blight. The input to that layer is 512 features learned by earlier layers. Suppose feature 87 has learned to fire on "dark concentric rings", a sign of early blight.

After training, the early-blight neuron has a large positive weight on feature 87, and the healthy neuron has a negative weight on it. The healthy neuron also has a larger bias, because 60% of the training photos were healthy, so its default score starts higher.

If the team retrains on a new dataset where only 20% of photos are healthy, the healthy neuron's bias drops. Same architecture, different numbers, different behaviour. That is exactly what "the knowledge is in the weights and biases" means.

Follow-up questions to expect

  • "Are biases regularised with weight decay?" — Usually not. Biases are few and do not cause overfitting in the same way, so many training setups exclude them, along with normalisation parameters.
  • "How many parameters does Linear(512, 10) have?" — 512 × 10 weights plus 10 biases, so 5,130.
  • "Can you remove bias if there is batch normalisation?" — Yes. Batch norm subtracts the mean and adds its own learned shift, so a bias in the layer right before it is redundant. That is why conv layers before BN often use bias=False.