Course Content
Deep Learning Essentials
13 sections · 61 lessons
Explain how a neural network works.
What you need to know
- Initialise — fill the weights with small random values (He or Xavier), biases usually with zeros.
- Forward pass — for a batch of inputs, each layer computes
activation(W x + b)until the output layer gives predictions. - Loss — compare predictions with the labels, giving one number for the batch.
- Backward pass — backpropagation computes the gradient of the loss with respect to every weight and bias.
- Update — the optimiser changes each parameter: in plain SGD,
w = w − learning_rate × gradient. - Repeat — for every batch in the dataset (one epoch), for many epochs, checking validation loss to know when to stop.
The whole loop in PyTorch
1import torch, torch.nn as nn2from torch.utils.data import TensorDataset, DataLoader34X = torch.randn(2000, 30) # 2,000 transactions, 30 features5y = (X[:, 0] + X[:, 1] * X[:, 2] > 1).float() # a non-linear "fraud" rule6loader = DataLoader(TensorDataset(X, y), batch_size=64, shuffle=True)78model = nn.Sequential(nn.Linear(30, 32), nn.ReLU(), nn.Linear(32, 1))9loss_fn = nn.BCEWithLogitsLoss()10opt = torch.optim.AdamW(model.parameters(), lr=1e-3)1112for epoch in range(20):13 for xb, yb in loader:14 logits = model(xb).squeeze(-1) # 1. forward pass15 loss = loss_fn(logits, yb) # 2. loss16 opt.zero_grad() # clear old gradients17 loss.backward() # 3. backward pass: gradients18 opt.step() # 4. update weightsFour lines inside the loop are the whole algorithm. zero_grad() is needed because PyTorch adds new gradients to old ones by default. With 2,000 rows and batch size 64, each epoch makes 32 updates, so 20 epochs make 640.
The intuition
At the start, the random weights give random predictions and a high loss. Each update is a small correction: weights that pushed towards wrong answers are reduced, and weights that helped are increased. After thousands of small corrections, hidden neurons have become detectors for useful patterns, and the output layer has learned how to combine them.
A real-life example
A speech-commands model trains on 85,000 one-second clips for 12 words. At epoch 0, it guesses almost uniformly, so the loss is about ln(12) = 2.48 and accuracy is about 8%, which is chance. After 1 epoch (about 1,300 updates with batch size 64), accuracy is 60%: early filters have learned to detect loud versus quiet and high versus low sounds. After 10 epochs, 93%: middle layers now respond to syllables. Validation accuracy stops improving at epoch 18, so training stops there and the epoch-18 weights are kept.
Follow-up questions to expect
- "What is an epoch?" — One full pass over the training data. An iteration is one batch update.
- "Why use mini-batches instead of the full dataset?" — Full-batch gradients are expensive and do not fit in memory. Mini-batches give a noisy but cheap estimate, and the noise even helps generalisation.
- "What happens if the learning rate is too high?" — The updates overshoot, the loss jumps around or becomes NaN. Too low, and training is very slow.