Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are batch, mini-batch, and stochastic gradient descent? Which one would you use and why?


What you need to know

Each variant estimates the same thing

The "true" gradient is the average gradient over every training example. All three methods estimate it; they differ in sample size.

  • Batch (full-batch) GD: average over all N examples, then update once. The estimate is exact, but you get one update per pass over the data.
  • Stochastic GD (SGD): compute the gradient on one example and update. Many updates per pass, each one a very rough estimate.
  • Mini-batch GD: average over B examples (the batch size), then update. The estimate's noise falls roughly with the square root of B, so a batch of 256 is about 16 times less noisy than a single example.

Why mini-batch wins on real hardware

A GPU processes 256 examples in about the same time as 1, because it does the maths in parallel. So SGD with one example leaves most of the GPU idle. Full batch, on the other hand, needs to hold activations for the whole dataset in memory, which is impossible for millions of images. Mini-batch sits in the sweet spot.

The noise in mini-batch gradients also has a benefit: it jostles the weights out of sharp, narrow minima, which tend to generalise worse than wide, flat ones.

Picking the batch size

  • Start with the largest power of 2 that fits comfortably in GPU memory, often 32 to 256 for images.
  • If you increase the batch size by a factor k, a common starting rule is to increase the learning rate by about k too, with a warmup. Very large batches without this adjustment often generalise worse.
  • If the batch you want does not fit, use gradient accumulation: run several small batches, add up their gradients, then update once.

In PyTorch, the batch size is set by the DataLoader, not by the optimiser:

Python
from torch.utils.data import DataLoaderloader = DataLoader(train_ds, batch_size=256, shuffle=True, num_workers=4)optimizer = torch.optim.SGD(model.parameters(), lr=0.05, momentum=0.9)

torch.optim.SGD is just the update rule. Fed with batches of 256, it performs mini-batch gradient descent.

MethodExamples per updateUpdates per epoch (1M rows)Gradient noiseGPU use
Batch GD1,000,0001NoneMemory usually runs out
Mini-batch GD2563,907ModerateGood
Stochastic GD11,000,000Very highPoor

A real-life example

A payments company trains a fraud model on 1 million labelled UPI transactions. Full-batch training would need the whole dataset's activations in GPU memory and would make only one update per epoch; after 10 epochs the model has taken just 10 steps. Pure SGD makes a million updates per epoch, but each GPU call processes one row, so an epoch takes hours and the loss curve looks like static.

With batches of 256, an epoch is 3,907 updates (the last batch is partial) and takes a few minutes. The team later tries batch size 4,096 to speed up training. Accuracy drops slightly until they scale the learning rate up and add a warmup, after which it matches the 256 run in a third of the time.

Follow-up questions to expect

  • "Why do people call mini-batch training 'SGD'?" — Historical habit. In deep learning, "SGD" almost always means mini-batch gradient descent, and PyTorch's optim.SGD is used with batches.
  • "Does a bigger batch always train faster?" — Each epoch runs faster, but beyond a point you need more epochs or a retuned learning rate to reach the same accuracy, and memory runs out.
  • "Why shuffle the data?" — So each batch is a fair sample. If data is sorted by class or by date, consecutive batches pull the weights in different directions and training becomes unstable.