Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between epoch, batch, and iteration?
What you need to know
The three terms
The link between them:
Text
iterations per epoch = ceil(dataset size / batch size)total iterations = iterations per epoch × number of epochsWith 50,000 images and batch size 64: 50,000 / 64 = 781.25, so 782 iterations, the last one a partial batch of 16. If you set drop_last=True in the DataLoader, the partial batch is skipped and you get 781.
Why the distinction matters in practice
- Learning-rate schedules. Some schedulers expect
scheduler.step()once per epoch (StepLRin its usual setup), others once per iteration (OneCycleLR). Calling a per-iteration scheduler once per epoch means the learning rate barely changes. - Logging and checkpointing. On huge datasets one epoch can take days, so teams count training in steps, not epochs. Large language models are often trained for less than one epoch over their data.
- Gradient accumulation. If you accumulate gradients over 4 batches of 64 before calling
optimizer.step(), the effective batch size is 256 and one optimiser step covers 4 forward-backward passes. Be clear whether "step" means a batch or an update.
A training loop makes the structure visible:
Python
1for epoch in range(10): # 10 epochs2 for xb, yb in loader: # 782 iterations per epoch3 loss = criterion(model(xb), yb) # one batch of 644 optimizer.zero_grad()5 loss.backward()6 optimizer.step() # one weight updateA real-life example
A small online furniture shop trains a product-photo classifier on 50,000 images with batch size 64 for 10 epochs, which is 7,820 updates. The engineer sets a cosine learning-rate schedule with T_max=10, thinking in epochs, but calls scheduler.step() inside the batch loop. The learning rate completes a full cosine cycle every 10 iterations, bouncing up and down about 780 times, and the loss curve is a saw-tooth. Setting T_max=7_820 fixes it.
Follow-up questions to expect
- "How many epochs should you train for?" — Until validation loss stops improving. Use early stopping rather than a fixed guess; small datasets with pretrained models often need 5 to 20 epochs, while training from scratch needs many more.
- "Is it bad if the last batch is smaller?" — Usually not. It can matter with batch normalisation if the last batch is very small, which is when
drop_last=Truehelps. - "What is an effective batch size?" — The number of examples that contribute to one optimiser update: per-device batch times number of devices times gradient-accumulation steps.