Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the difference between epoch, batch, and iteration?


What you need to know

The three terms

The link between them:

Text
iterations per epoch = ceil(dataset size / batch size)total iterations     = iterations per epoch × number of epochs

With 50,000 images and batch size 64: 50,000 / 64 = 781.25, so 782 iterations, the last one a partial batch of 16. If you set drop_last=True in the DataLoader, the partial batch is skipped and you get 781.

Why the distinction matters in practice

  • Learning-rate schedules. Some schedulers expect scheduler.step() once per epoch (StepLR in its usual setup), others once per iteration (OneCycleLR). Calling a per-iteration scheduler once per epoch means the learning rate barely changes.
  • Logging and checkpointing. On huge datasets one epoch can take days, so teams count training in steps, not epochs. Large language models are often trained for less than one epoch over their data.
  • Gradient accumulation. If you accumulate gradients over 4 batches of 64 before calling optimizer.step(), the effective batch size is 256 and one optimiser step covers 4 forward-backward passes. Be clear whether "step" means a batch or an update.

A training loop makes the structure visible:

Python
for epoch in range(10):                  # 10 epochs    for xb, yb in loader:                # 782 iterations per epoch        loss = criterion(model(xb), yb)  # one batch of 64        optimizer.zero_grad()        loss.backward()        optimizer.step()                 # one weight update

A real-life example

A small online furniture shop trains a product-photo classifier on 50,000 images with batch size 64 for 10 epochs, which is 7,820 updates. The engineer sets a cosine learning-rate schedule with T_max=10, thinking in epochs, but calls scheduler.step() inside the batch loop. The learning rate completes a full cosine cycle every 10 iterations, bouncing up and down about 780 times, and the loss curve is a saw-tooth. Setting T_max=7_820 fixes it.

Follow-up questions to expect

  • "How many epochs should you train for?" — Until validation loss stops improving. Use early stopping rather than a fixed guess; small datasets with pretrained models often need 5 to 20 epochs, while training from scratch needs many more.
  • "Is it bad if the last batch is smaller?" — Usually not. It can matter with batch normalisation if the last batch is very small, which is when drop_last=True helps.
  • "What is an effective batch size?" — The number of examples that contribute to one optimiser update: per-device batch times number of devices times gradient-accumulation steps.