Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is data or image augmentation?
What you need to know
What happens to the data
Augmentation does not usually save extra copies to disk. Each time the data loader fetches an image, it applies random transforms, so in epoch 1 the model might see a shoe cropped slightly left and brightened, and in epoch 2 the same shoe flipped and a little darker. Over 30 epochs, the model sees 30 different versions of every photo.
Common augmentations by data type
- Images: random resized crop, horizontal flip, small rotation, colour and brightness jitter, blur, cutout (blank a random patch). Stronger mixing methods, MixUp (blend two images and their labels) and CutMix (paste a patch of one image onto another, mixing labels by area), are common in modern recipes.
- Text: synonym swaps, random deletion, back-translation (translate to Hindi and back to English), and paraphrases generated by an LLM.
- Audio: time shift, speed and pitch change, added background noise, and SpecAugment, which masks bands of time and frequency in the spectrogram.
- Tabular: rarely used; small noise on continuous features at most.
A current torchvision pipeline
1import torch2from torchvision.transforms import v234train_tf = v2.Compose([5 v2.RandomResizedCrop(224, scale=(0.7, 1.0)),6 v2.RandomHorizontalFlip(p=0.5),7 v2.ColorJitter(brightness=0.2, contrast=0.2),8 v2.ToImage(), v2.ToDtype(torch.float32, scale=True),9 v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),10])11val_tf = v2.Compose([12 v2.Resize(256), v2.CenterCrop(224),13 v2.ToImage(), v2.ToDtype(torch.float32, scale=True),14 v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),15])Training images get random crops, flips and colour changes. Validation images get only a fixed resize and centre crop, so the evaluation is the same every time. torchvision.transforms.v2 is the current API; the older transforms module still works.
Keep it label-preserving
Ask: "Would a human still give this image the same label?"
- A vertically flipped "6" becomes a "9".
- Mirrored product photos reverse printed text and logos.
- A horizontally flipped chest X-ray puts the heart on the wrong side. It is anatomically unrealistic, so many teams skip flips for X-rays and use small rotations and contrast changes instead.
- Heavy colour jitter can destroy the signal when colour is the label, such as ripe versus unripe fruit.
A real-life example
A small home-decor brand wants to auto-categorise product photos into 12 categories but has only 1,800 labelled images, all shot in a studio on a white background. Sellers on its marketplace upload phone photos: tilted, in warm indoor light, with the product off-centre.
Trained without augmentation, a pretrained ResNet reaches 97% on the studio validation set and 68% on 300 real seller photos. The team adds random resized crops (so off-centre products are normal), small rotations up to 15 degrees, strong brightness and colour-temperature jitter (for indoor lighting), and horizontal flips. They skip flips for the "wall art" category, where mirrored text looks wrong. Accuracy on seller photos rises to 86%, with no new labelling.
Follow-up questions to expect
- "Should you augment the validation or test set?" — No. They must reflect real data so the metric is honest. The exception is test-time augmentation, where you deliberately average predictions over a few flips or crops at inference, and you evaluate with the same procedure.
- "Does augmentation increase the dataset size?" — Not on disk. It increases the variety the model sees across epochs.
- "Can augmentation hurt?" — Yes, if transforms change the label or are much stronger than anything in real data. It can also slow convergence, so heavily augmented models often need more epochs.