Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are batch, mini-batch, and stochastic gradient descent? Which one would you use and why?
What you need to know
Each variant estimates the same thing
The "true" gradient is the average gradient over every training example. All three methods estimate it; they differ in sample size.
- Batch (full-batch) GD: average over all N examples, then update once. The estimate is exact, but you get one update per pass over the data.
- Stochastic GD (SGD): compute the gradient on one example and update. Many updates per pass, each one a very rough estimate.
- Mini-batch GD: average over B examples (the batch size), then update. The estimate's noise falls roughly with the square root of B, so a batch of 256 is about 16 times less noisy than a single example.
Why mini-batch wins on real hardware
A GPU processes 256 examples in about the same time as 1, because it does the maths in parallel. So SGD with one example leaves most of the GPU idle. Full batch, on the other hand, needs to hold activations for the whole dataset in memory, which is impossible for millions of images. Mini-batch sits in the sweet spot.
The noise in mini-batch gradients also has a benefit: it jostles the weights out of sharp, narrow minima, which tend to generalise worse than wide, flat ones.
Picking the batch size
- Start with the largest power of 2 that fits comfortably in GPU memory, often 32 to 256 for images.
- If you increase the batch size by a factor k, a common starting rule is to increase the learning rate by about k too, with a warmup. Very large batches without this adjustment often generalise worse.
- If the batch you want does not fit, use gradient accumulation: run several small batches, add up their gradients, then update once.
In PyTorch, the batch size is set by the DataLoader, not by the optimiser:
1from torch.utils.data import DataLoader23loader = DataLoader(train_ds, batch_size=256, shuffle=True, num_workers=4)4optimizer = torch.optim.SGD(model.parameters(), lr=0.05, momentum=0.9)torch.optim.SGD is just the update rule. Fed with batches of 256, it performs mini-batch gradient descent.
| Method | Examples per update | Updates per epoch (1M rows) | Gradient noise | GPU use |
|---|---|---|---|---|
| Batch GD | 1,000,000 | 1 | None | Memory usually runs out |
| Mini-batch GD | 256 | 3,907 | Moderate | Good |
| Stochastic GD | 1 | 1,000,000 | Very high | Poor |
A real-life example
A payments company trains a fraud model on 1 million labelled UPI transactions. Full-batch training would need the whole dataset's activations in GPU memory and would make only one update per epoch; after 10 epochs the model has taken just 10 steps. Pure SGD makes a million updates per epoch, but each GPU call processes one row, so an epoch takes hours and the loss curve looks like static.
With batches of 256, an epoch is 3,907 updates (the last batch is partial) and takes a few minutes. The team later tries batch size 4,096 to speed up training. Accuracy drops slightly until they scale the learning rate up and add a warmup, after which it matches the 256 run in a third of the time.
Follow-up questions to expect
- "Why do people call mini-batch training 'SGD'?" — Historical habit. In deep learning, "SGD" almost always means mini-batch gradient descent, and PyTorch's
optim.SGDis used with batches. - "Does a bigger batch always train faster?" — Each epoch runs faster, but beyond a point you need more epochs or a retuned learning rate to reach the same accuracy, and memory runs out.
- "Why shuffle the data?" — So each batch is a fair sample. If data is sorted by class or by date, consecutive batches pull the weights in different directions and training becomes unstable.