Course Content
Deep Learning Essentials
13 sections · 61 lessons
What happens during the forward pass in a neural network?
What you need to know
The computation for one layer
z = W x + b (pre-activation: a linear transformation)a = activation(z) (the layer's output, which feeds the next layer)For a batch, x is a matrix with one row per example, so one matrix multiply handles the whole batch at once.
A forward pass by hand
A tiny network: 2 inputs, 2 hidden neurons with ReLU, 1 output with sigmoid.
x = [1.0, 2.0]W1 = [[0.5, -0.2], b1 = [0.1, -0.5] [0.3, 0.8]]z1 = W1 x + b1 = [0.5 - 0.4 + 0.1, 0.3 + 1.6 - 0.5] = [0.2, 1.4]a1 = ReLU(z1) = [0.2, 1.4]W2 = [1.0, -0.5], b2 = 0.2z2 = 1.0*0.2 - 0.5*1.4 + 0.2 = -0.3 (the logit)p = sigmoid(-0.3) = 0.426 (probability of class 1)If the true label is 1, the binary cross-entropy loss is -ln(0.426) = 0.854. The backward pass will then work out how to change each of the nine parameters to raise p.
The same thing in PyTorch
1import torch, torch.nn as nn23model = nn.Sequential(4 nn.Flatten(), # (8, 1, 40, 101) -> (8, 4040)5 nn.Linear(40 * 101, 64), nn.ReLU(),6 nn.Linear(64, 12), # 12 spoken commands7)8x = torch.randn(8, 1, 40, 101) # batch of 8 spectrograms910logits = model(x) # forward pass: shape (8, 12)11probs = logits.softmax(dim=-1) # each row sums to 1Calling model(x) runs each layer's forward in order. The shape changes are the story: 8 spectrograms become 8 rows of 64 features, then 8 rows of 12 scores.
Training mode vs inference
- During training, PyTorch records the graph and stores intermediate activations (
z1,a1above) because the gradient of each weight depends on them. For large models, these stored activations use more memory than the weights. That is why batch size is limited by GPU memory. - At inference, use
model.eval()so dropout and batch norm behave correctly, andtorch.inference_mode()(ortorch.no_grad()) so no graph or activations are kept. This saves memory and time.
model.eval()with torch.inference_mode(): pred = model(x).argmax(dim=-1)A real-life example
A smart speaker listens for 12 commands like "yes", "no", "stop", "go". Every second, the device turns audio into a 40×101 spectrogram and runs one forward pass of a small CNN:
- Conv layers turn the spectrogram into feature maps that respond to short sound patterns.
- Pooling shrinks them; a final layer produces 12 logits.
- Softmax converts them to probabilities, for example "stop" 0.91, "go" 0.04, others less.
- If the top probability is above 0.8, the device acts.
On the device, only the forward pass runs, in about 5 milliseconds, with no gradients and no stored activations. During training on a server, the same forward pass runs on batches of 256 clips, keeps every activation, and is followed by a backward pass that roughly doubles the compute.
Follow-up questions to expect
- "Why keep activations during the forward pass?" — The gradient for a weight depends on the input it saw. For
z = W a + b, the gradient ofWneedsa, so it must be stored until the backward pass. - "What is the difference between
no_gradandinference_mode?" — Both stop graph recording.inference_modealso skips some extra bookkeeping, so it is slightly faster, but its tensors cannot be used later in autograd. - "What is gradient checkpointing?" — Storing only some activations and recomputing the rest during the backward pass. It trades extra compute for much less memory.