Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Explain how a neural network works.


The training looprandomweightsforward passlossbackwardpassupdateweightsgradient per weightw minus lrtimes grad2,000 rows in batches of 64 is 32 updates per epoch.
Learning is thousands of small corrections: each loop nudges every weight a little in the direction that lowers the loss.

What you need to know

  1. Initialise — fill the weights with small random values (He or Xavier), biases usually with zeros.
  2. Forward pass — for a batch of inputs, each layer computes activation(W x + b) until the output layer gives predictions.
  3. Loss — compare predictions with the labels, giving one number for the batch.
  4. Backward pass — backpropagation computes the gradient of the loss with respect to every weight and bias.
  5. Update — the optimiser changes each parameter: in plain SGD, w = w − learning_rate × gradient.
  6. Repeat — for every batch in the dataset (one epoch), for many epochs, checking validation loss to know when to stop.

The whole loop in PyTorch

Python
import torch, torch.nn as nnfrom torch.utils.data import TensorDataset, DataLoaderX = torch.randn(2000, 30)                         # 2,000 transactions, 30 featuresy = (X[:, 0] + X[:, 1] * X[:, 2] > 1).float()     # a non-linear "fraud" ruleloader = DataLoader(TensorDataset(X, y), batch_size=64, shuffle=True)model = nn.Sequential(nn.Linear(30, 32), nn.ReLU(), nn.Linear(32, 1))loss_fn = nn.BCEWithLogitsLoss()opt = torch.optim.AdamW(model.parameters(), lr=1e-3)for epoch in range(20):    for xb, yb in loader:        logits = model(xb).squeeze(-1)    # 1. forward pass        loss = loss_fn(logits, yb)        # 2. loss        opt.zero_grad()                   #    clear old gradients        loss.backward()                   # 3. backward pass: gradients        opt.step()                        # 4. update weights

Four lines inside the loop are the whole algorithm. zero_grad() is needed because PyTorch adds new gradients to old ones by default. With 2,000 rows and batch size 64, each epoch makes 32 updates, so 20 epochs make 640.

The intuition

At the start, the random weights give random predictions and a high loss. Each update is a small correction: weights that pushed towards wrong answers are reduced, and weights that helped are increased. After thousands of small corrections, hidden neurons have become detectors for useful patterns, and the output layer has learned how to combine them.

A real-life example

A speech-commands model trains on 85,000 one-second clips for 12 words. At epoch 0, it guesses almost uniformly, so the loss is about ln(12) = 2.48 and accuracy is about 8%, which is chance. After 1 epoch (about 1,300 updates with batch size 64), accuracy is 60%: early filters have learned to detect loud versus quiet and high versus low sounds. After 10 epochs, 93%: middle layers now respond to syllables. Validation accuracy stops improving at epoch 18, so training stops there and the epoch-18 weights are kept.

Follow-up questions to expect

  • "What is an epoch?" — One full pass over the training data. An iteration is one batch update.
  • "Why use mini-batches instead of the full dataset?" — Full-batch gradients are expensive and do not fit in memory. Mini-batches give a noisy but cheap estimate, and the noise even helps generalisation.
  • "What happens if the learning rate is too high?" — The updates overshoot, the loss jumps around or becomes NaN. Too low, and training is very slow.