Computer Vision Fundamentals

Data Augmentation & Normalization


You have 800 photographs of skin lesions, labelled benign or malignant by a dermatologist. That is a small dataset by modern standards but a genuinely expensive one — each label cost expert time. You train a reasonable convolutional network on it.

Training accuracy climbs to 99.6%. Validation accuracy stalls at 71% and then starts falling. The gap widens with every epoch.

The network has not learned dermatology. It has learned those 800 specific photographs. There is a mole in image 214 with a particular arrangement of hairs beside it; the network has memorised that arrangement. There is a ruler in the corner of forty images, all of which happen to be malignant; the network has learned to look for rulers. None of this generalises, because none of it is about the lesion.

The obvious fix is more data. But the dermatologist is busy and each label costs real money. So here is the question that data augmentation answers: can you manufacture more training examples from the ones you already have?

The families you can manufacture data from800 lesion imagesSpatial: flip, rotate, cropColour:brightness, contrast, hueNoise, blur and occlusionMixing: mixup and CutMixTest-time augmentation
Augmentation does not add information; it tells the model which changes must not change the label.

Why manufactured data helps at all

The instinct is that it should not. Rotating a photograph adds no new information — the pixels are a deterministic function of pixels you already had. How can rearranging what you have teach the model something it did not know?

The resolution is that augmentation does not add information about the subject. It adds information about what is irrelevant.

Show the network the same lesion at 0°, 15° and 340°, all labelled malignant, and you have told it something it could not otherwise know: rotation does not change the diagnosis. That is a real constraint, and it was not in the original data. Every augmentation you apply is a statement of the form "this variation should not change the answer" — and the model learns to be invariant to it.

Augmentation does not create data. It encodes your knowledge of which variations are meaningless, and forces the model to stop relying on them.

This also explains why augmentation fights the specific failure above. If the hairs beside mole 214 appear sometimes flipped, sometimes shifted, sometimes with different brightness, they stop being a reliable fingerprint for that one image. Memorisation gets harder; the actual lesion features, which survive every transformation, get comparatively easier. You have made the shortcut more expensive than the real solution.

The effective dataset size

With 800 images and a pipeline that applies random rotation, flip, crop and colour jitter, the network essentially never sees the same array twice. If each augmentation has a continuous parameter, the number of distinct variants is unbounded in principle. In practice the useful multiplier is modest — the variants are highly correlated with each other — but the effect on overfitting is real and large.

The main families of transformation

Spatial

TransformTypical rangeEncodes the belief that…Watch out for
Horizontal flipp = 0.5Left and right are interchangeableFalse for text, digits, and left/right anatomy
Vertical flipp = 0.5Up and down are interchangeableAlmost always false for photographs of the world
Rotation±10° to ±30°Camera tilt is incidentalCorners get cut or filled; large angles create artificial borders
Random crop / scale60–100% of areaFraming and distance are incidentalCan crop the subject out entirely
Translation±10% of sizePosition in frame does not matterLeaves empty strips
Shear±10°Mild viewing-angle distortion is incidentalUnnatural at large values

Random resized crop is the workhorse for natural images and does more than it appears to. It varies position, scale and aspect ratio at once, and it forces the network to recognise objects from partial views — which is exactly what happens at inference when something is half occluded.

For satellite and microscopy imagery, both flips and full 90° rotations are fair game, because there is no canonical "up". For street scenes, vertical flipping is nonsense: the model would learn that upside-down cars exist, and waste capacity on a class of input it will never encounter.

Colour and brightness

These simulate different cameras, sensors and lighting rather than different geometry.

  • Brightness — add or scale intensity, typically ±20–40%. Simulates exposure differences.
  • Contrast — stretch or compress values around the mean. Simulates flat versus punchy lighting.
  • Saturation — scale distance from grey. Simulates camera colour processing.
  • Hue shift — rotate the colour wheel, usually only a few degrees. Larger shifts turn grass purple.
  • Grayscale conversion — with low probability, drops colour entirely, forcing reliance on shape and texture.

Colour augmentation carries the highest risk of destroying the label, and the risk depends entirely on the task. Hue shifts on ImageNet photographs are harmless. Hue shifts on traffic-light images are catastrophic: shift red far enough and it becomes green, and you have handed the model an image labelled "stop" that shows a green light. The model does not know your augmentation was a mistake. It learns the contradiction and gets worse.

The same logic bans aggressive colour jitter in medical imaging where colour encodes tissue type, in agricultural sorting where colour indicates ripeness, and in quality inspection where discolouration is the defect.

Noise, blur and occlusion

Gaussian noise simulates low-light sensor grain. Motion blur simulates a moving camera or subject. Both make the model robust to poor input quality, which matters enormously if it will run on cheap cameras or handheld phones.

Occlusion augmentation is worth calling out separately. Random erasing (also called cutout) blanks a random rectangle of the image. It sounds destructive, and it works well, because it prevents the model from depending on any one region. If the network has learned to identify a dog purely from its ears, erasing the ears half the time forces it to also learn the body, the legs and the texture of the fur. The result is a model that still works when the ears are behind a fence.

Mixing images together

Two techniques go further and combine multiple training images:

TechniqueImage operationLabel operation
MixupPixel-wise blend: λx1+(1−λ)x2\lambda x_1 + (1-\lambda) x_2Blend labels the same way
CutMixPaste a rectangle from image 2 into image 1Weight labels by the area each occupies

A mixup image is 70% cat and 30% dog, and its label is 0.7 cat, 0.3 dog. This looks bizarre — no such photograph exists in the world — but it produces measurably better-calibrated models, because it forces predictions to vary smoothly between classes instead of jumping from total confidence in one to total confidence in another.

Choosing how hard to push

Augmentation strength is a hyperparameter, and both extremes fail in ways you can diagnose from the training curves.

SettingSymptomReading of the curves
Too weakOverfittingTraining accuracy near 100%, validation far below and drifting down
About rightHealthyTraining slightly above validation, both still improving
Too strongUnderfittingTraining accuracy stuck low; training loss barely below validation loss

That last row is the one people misread. When training accuracy will not rise, the reflex is to add capacity or train longer. But if you have applied ±45° rotation, heavy colour jitter, strong blur and aggressive cutout to a set of 800 images, you have made the training task genuinely harder than the real task. The model is not failing to learn; it is correctly learning an impossible problem.

Augmentation strength is a dial, not a checklist. Too little and the model memorises the training set; too much and you have set it an exam harder than the one it will actually sit.

Sensible starting points, then adjust based on the gap between training and validation:

SituationStrengthTypical set
Large dataset (>100k), big modelLightFlip, small random crop
Medium dataset (10k–100k)ModerateFlip, crop, ±15° rotation, mild colour jitter
Small dataset (<10k)HeavyAll of the above plus cutout, blur, mixup
Known deployment shift (different cameras, lighting)TargetedWhatever specifically simulates the shift

That last row is the most valuable and the most neglected. If your model will run on a factory camera under fluorescent light while your training photos were taken under daylight, the single highest-value augmentation is a colour temperature shift. Generic augmentation lists are a starting point; the real question is always "what will actually be different when this runs for real?"

Where augmentation belongs in the pipeline

Two rules, both of which people break.

Augment the training set only. Validation and test data must be processed exactly as production data will be — resize, normalise, nothing else. If you augment validation, its score measures performance on distorted images that will never occur, and it fluctuates randomly between epochs, so you cannot tell whether a change helped. Worse, model selection based on a noisy metric picks essentially at random.

Augment on the fly, not on disk. Pre-generating ten augmented copies of each image and saving them gives a fixed set of ten variants that the model sees repeatedly and can memorise. Applying transforms in the data loader, with fresh random parameters each time, means the model genuinely never sees the same array twice. It also costs no storage.

The one deliberate exception: test-time augmentation

Test-time augmentation is the exception that proves the rule, and it is doing something different. At inference, run several augmented versions of the same input through the model and average the predictions.

Python
import torch@torch.no_grad()def predict_tta(model, image):    views = [image, torch.flip(image, dims=[-1])]   # original + horizontal flip    probs = [torch.softmax(model(v.unsqueeze(0)), dim=1) for v in views]    return torch.stack(probs).mean(dim=0)

This typically buys 0.5–2 percentage points of accuracy for a proportional increase in inference cost. It works because the errors made on different views are partly independent, so averaging cancels some of them. Use only label-preserving transforms — a horizontal flip and a couple of crops is the standard set. It is a good trade for offline batch scoring and usually a bad one for real-time systems, where doubling latency matters more than a point of accuracy.

Normalisation alongside augmentation

Normalisation is a separate concern from augmentation and follows a different rule: identical everywhere. Augmentation differs between training and inference by design. Normalisation must not differ at all.

The two common schemes:

SchemeFormulaOutput rangeUse when
Scale to unit intervalx/255x / 255[0,1][0, 1]Training from scratch; simple pipelines
Standardise(x−μ)/σ(x - \mu) / \sigmaroughly [−3,3][-3, 3]Pretrained backbones; deeper networks

For an ImageNet-pretrained backbone, use the ImageNet statistics — mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225], per channel, after scaling to [0,1][0, 1]. Those filters were tuned for inputs with that distribution. Feeding raw [0,1][0, 1] values instead does not crash anything; it just quietly costs accuracy, and the cause is nearly invisible when you go looking.

For a model trained from scratch, compute the mean and standard deviation over your training split only. Including validation data leaks distributional information and inflates your validation score by a small, unquantifiable amount.

Order within the pipeline

Augment first, normalise last. The reason is mechanical: most augmentation libraries expect uint8 values in 0–255, because that is what brightness, contrast and hue adjustments are defined against. Hand them standardised floats with negative values and brightness adjustment either errors out or produces nonsense. Convert to tensor after the pixel-level work is done, then normalise.

Python
from torchvision.transforms import v2import torchMEAN, STD = [0.485, 0.456, 0.406], [0.229, 0.224, 0.225]train_tf = v2.Compose([    v2.RandomResizedCrop(224, scale=(0.7, 1.0)),    v2.RandomHorizontalFlip(p=0.5),    v2.RandomRotation(degrees=15),    v2.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2, hue=0.02),    v2.ToImage(),    v2.ToDtype(torch.float32, scale=True),   # -> [0, 1]    v2.Normalize(mean=MEAN, std=STD),    v2.RandomErasing(p=0.25),                # operates on the tensor, so after])eval_tf = v2.Compose([    v2.Resize(256),    v2.CenterCrop(224),    v2.ToImage(),    v2.ToDtype(torch.float32, scale=True),    v2.Normalize(mean=MEAN, std=STD),])

Note that RandomErasing sits after normalisation. It blanks regions with the dataset mean, which in normalised space is zero, so it belongs on the tensor side. That is a deliberate exception, not a contradiction.

Applying this to your own problem

Before adding any transform, apply one test: would a human expert still give this image the same label? If a dermatologist would still call the rotated, flipped, slightly darker lesion malignant, the augmentation is valid. If a hue shift turns a red traffic light green, it is not — you have created a mislabelled example, and mislabelled examples actively damage a model rather than merely failing to help.

Then look at the output. Write a loop that pulls a batch from the augmented training loader, denormalises it, and saves a grid as an image. Two minutes of looking catches upside-down photographs, subjects cropped clean out of frame, colour shifts that changed the class, and masks that were rotated while their images were not. None of those show up in a loss curve, and all of them are obvious to your eye.

Finally, treat augmentation as a response to a diagnosis rather than a default. Wide train-validation gap means overfitting, so increase strength. Training accuracy that refuses to climb means the task has been made too hard, so reduce it. A model that works in the lab and fails in the field means a distribution shift you have not simulated — so find out precisely what is different out there, and add the transform that reproduces it.