Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What happens during the forward pass in a neural network?


A forward pass through a 2-2-1 networkx = [1.0, 2.0]z1 = W1 x + b1= [0.2, 1.4]ReLU: a1 =[0.2, 1.4]z2 = W2 a1+ b2 = −0.3sigmoid:p = 0.426Loss if y= 1: 0.854Training keeps z1 and a1 in memory; inference throws them away.
Every intermediate number is stored during training because the gradient of each weight depends on the input that weight saw.

What you need to know

The computation for one layer

Text
z = W x + b          (pre-activation: a linear transformation)a = activation(z)    (the layer's output, which feeds the next layer)

For a batch, x is a matrix with one row per example, so one matrix multiply handles the whole batch at once.

A forward pass by hand

A tiny network: 2 inputs, 2 hidden neurons with ReLU, 1 output with sigmoid.

Text
x  = [1.0, 2.0]W1 = [[0.5, -0.2],     b1 = [0.1, -0.5]      [0.3,  0.8]]z1 = W1 x + b1  = [0.5 - 0.4 + 0.1,  0.3 + 1.6 - 0.5] = [0.2, 1.4]a1 = ReLU(z1)   = [0.2, 1.4]W2 = [1.0, -0.5],  b2 = 0.2z2 = 1.0*0.2 - 0.5*1.4 + 0.2 = -0.3        (the logit)p  = sigmoid(-0.3) = 0.426                 (probability of class 1)

If the true label is 1, the binary cross-entropy loss is -ln(0.426) = 0.854. The backward pass will then work out how to change each of the nine parameters to raise p.

The same thing in PyTorch

Python
import torch, torch.nn as nnmodel = nn.Sequential(    nn.Flatten(),                      # (8, 1, 40, 101) -> (8, 4040)    nn.Linear(40 * 101, 64), nn.ReLU(),    nn.Linear(64, 12),                 # 12 spoken commands)x = torch.randn(8, 1, 40, 101)         # batch of 8 spectrogramslogits = model(x)                      # forward pass: shape (8, 12)probs = logits.softmax(dim=-1)         # each row sums to 1

Calling model(x) runs each layer's forward in order. The shape changes are the story: 8 spectrograms become 8 rows of 64 features, then 8 rows of 12 scores.

Training mode vs inference

  • During training, PyTorch records the graph and stores intermediate activations (z1, a1 above) because the gradient of each weight depends on them. For large models, these stored activations use more memory than the weights. That is why batch size is limited by GPU memory.
  • At inference, use model.eval() so dropout and batch norm behave correctly, and torch.inference_mode() (or torch.no_grad()) so no graph or activations are kept. This saves memory and time.
Python
model.eval()with torch.inference_mode():    pred = model(x).argmax(dim=-1)

A real-life example

A smart speaker listens for 12 commands like "yes", "no", "stop", "go". Every second, the device turns audio into a 40×101 spectrogram and runs one forward pass of a small CNN:

  1. Conv layers turn the spectrogram into feature maps that respond to short sound patterns.
  2. Pooling shrinks them; a final layer produces 12 logits.
  3. Softmax converts them to probabilities, for example "stop" 0.91, "go" 0.04, others less.
  4. If the top probability is above 0.8, the device acts.

On the device, only the forward pass runs, in about 5 milliseconds, with no gradients and no stored activations. During training on a server, the same forward pass runs on batches of 256 clips, keeps every activation, and is followed by a backward pass that roughly doubles the compute.

Follow-up questions to expect

  • "Why keep activations during the forward pass?" — The gradient for a weight depends on the input it saw. For z = W a + b, the gradient of W needs a, so it must be stored until the backward pass.
  • "What is the difference between no_grad and inference_mode?" — Both stop graph recording. inference_mode also skips some extra bookkeeping, so it is slightly faster, but its tensors cannot be used later in autograd.
  • "What is gradient checkpointing?" — Storing only some activations and recomputing the rest during the backward pass. It trades extra compute for much less memory.