Course Content
Computer Vision Fundamentals
3 sections · 8 lessons
Image Preprocessing Techniques
A team trains a classifier to spot manufacturing defects on a production line. Validation accuracy: 97%. They deploy it. In the factory it gets about 60% — barely better than guessing between two classes.
Nothing is wrong with the model. The training images came from a folder of 4000×3000 photographs, resized by a script that squashed everything to 224×224 regardless of shape. The live camera produces 1280×720 frames, which squash to a different aspect ratio entirely. A circular defect that appeared as a circle in training arrives at inference as an ellipse. The model has never seen an ellipse.
That is a preprocessing bug, and preprocessing bugs have a signature: the model looks fine on your data and falls apart on anyone else's. They are also, unlike architecture problems, entirely preventable — provided you understand what each transformation actually does to the numbers rather than treating the pipeline as a magic incantation copied from a tutorial.
Preprocessing is everything that happens between reading a file from disk and handing an array to a model. Its job is to make images consistent in size, consistent in numeric range, and free of variation that has nothing to do with the task.
Resizing, and why it is never free
Neural networks with fully connected layers require a fixed input size, and batching requires every image in a batch to have identical dimensions. So resizing is almost always the first step. The question is how to invent pixels that were never measured.
Interpolation: guessing the values in between
Enlarging a 2×2 image to 4×4 means producing 16 values from 4. Twelve of them do not exist and must be estimated. That estimation is interpolation, and the method you choose is a trade between speed and smoothness.
| Method | How it decides a value | Speed | Result | Use for |
|---|---|---|---|---|
| Nearest neighbour | Copies the closest original pixel | Fastest | Blocky, hard edges | Segmentation masks and label maps |
| Bilinear | Weighted average of the 4 surrounding pixels | Fast | Smooth, slightly soft | General-purpose default |
| Bicubic | Cubic fit over the 16 surrounding pixels | Slower | Sharper than bilinear | Upscaling where quality matters |
| Lanczos | Windowed sinc over a wider neighbourhood | Slowest | Sharpest, can ring at edges | Final-quality photographic resizing |
| Area | Averages all pixels falling into the output cell | Fast | Clean, no aliasing | Downscaling |
Two of those rows are not preferences but rules, and both have a concrete failure attached.
Use nearest neighbour for masks. A segmentation mask holds class identifiers: 0 background, 1 road, 2 car. Bilinear interpolation averages neighbouring values, so a pixel on the boundary between class 0 and class 2 becomes 1 — a road pixel that exists nowhere in the original annotation. Your labels now contain fabricated classes along every object edge. Nearest neighbour copies rather than blends, so identifiers stay identifiers.
Use area for downscaling. Shrinking 1000×1000 to 100×100 with bilinear interpolation samples roughly one pixel out of every hundred and discards the rest. If the image contains a fine repeating pattern — a striped shirt, a mesh grille, text — the sampling grid beats against the pattern and produces aliasing: moiré fringes and stripes that were never there. Area interpolation averages over the whole 10×10 block that maps to each output pixel, so all the information contributes and the artefacts vanish.
Downscale by averaging, upscale by interpolating, and resize masks by copying. Almost every resizing artefact you will ever see comes from breaking one of those three.
Aspect ratio: squash, crop, or pad
A 1600×900 image needs to become 224×224. Three options, three different costs:
| Strategy | What happens | Cost | Sensible when |
|---|---|---|---|
| Stretch to fit | Both axes scaled independently | Geometry is distorted; circles become ellipses | Shape genuinely does not matter (texture, some medical modalities) |
| Centre crop | Scale the short side to 224, cut the long side | Content at the edges is thrown away | The subject is reliably central |
| Letterbox / pad | Scale so the long side is 224, fill the rest | Wasted pixels; the model must learn to ignore the padding | Detection, where every object must survive and boxes must stay valid |
The defect-detection failure at the top of this page was a stretch applied inconsistently between training and deployment. Note that stretching was not the real sin — using different aspect ratios in training and production was. Consistency matters more than which option you pick.
Letterboxing is the standard choice for object detection because cropping can remove an object entirely, and stretching invalidates the aspect ratio priors that anchor-based detectors rely on. The padding is usually a neutral grey (114, 114, 114) rather than black, so the border does not look like a strong dark edge to the first convolution layer.
1import cv22import numpy as np34def letterbox(img, size=224, fill=114):5 h, w = img.shape[:2]6 scale = size / max(h, w)7 nh, nw = round(h * scale), round(w * scale)8 # shrinking -> AREA, growing -> LINEAR9 interp = cv2.INTER_AREA if scale < 1 else cv2.INTER_LINEAR10 resized = cv2.resize(img, (nw, nh), interpolation=interp)1112 canvas = np.full((size, size, 3), fill, dtype=np.uint8)13 top, left = (size - nh) // 2, (size - nw) // 214 canvas[top:top + nh, left:left + nw] = resized15 return canvas, scale, (left, top) # keep scale/offset to map boxes backReturning the scale and offset is not optional housekeeping. Without them you cannot convert a predicted box in 224-space back to a box in the original image, and you will end up with detections drawn in the wrong place.
Getting the numbers into a sensible range
Raw pixel values run 0 to 255. Feeding those directly to a network trains badly, and it is worth understanding exactly why rather than accepting "you should normalise" as folklore.
The first layer computes a weighted sum of its inputs. With inputs around 150 and a few hundred of them, that sum is in the tens of thousands before any weight is learned. Push that through a sigmoid or tanh and you land deep in the flat region where the gradient is effectively zero — the layer stops learning. Even with ReLU, which does not saturate, the gradients with respect to the weights are proportional to the input magnitude, so large inputs give large gradients, and the learning rate that keeps the first layer stable is far too small for the rest of the network.
There is a second, subtler problem. If one input dimension typically sits near 200 and another near 10, the loss surface becomes a long narrow valley, and gradient descent zig-zags across it instead of running down it. Equalising the scales makes the valley round and the descent direct.
Large inputs do not merely slow training down. They force a learning rate small enough to keep the first layer stable, and that rate is far too small for every layer after it.
Min-max scaling
For 8-bit images this is just division by 255, mapping everything to [0,1]. Simple, bounded, and preserves the shape of the distribution exactly. Its weakness is sensitivity to outliers: if a single hot pixel reads 255 in an otherwise dim 16-bit scan, that one pixel sets the maximum and compresses everything real into the bottom of the range.
Standardisation (z-score)
Subtract the mean, divide by the standard deviation. The result has mean 0 and standard deviation 1, and is not bounded — values typically land in roughly [−3,3].
Centring on zero matters more than it looks. If every input to a neuron is positive, every weight's gradient shares the same sign on a given step, so all the weights of that neuron must move up together or down together. Progress happens in zig-zags. Zero-centred inputs let individual weights move in whichever direction they need to.
The critical practical point is which mean and standard deviation you use. For ImageNet-pretrained models, use the ImageNet statistics — the network's learned filters expect inputs distributed that way:
IMAGENET_MEAN = [0.485, 0.456, 0.406]IMAGENET_STD = [0.229, 0.224, 0.225]For a model trained from scratch on your own data, compute the statistics from your training set only. Computing them over train plus validation leaks information about the validation set into training, and your validation score becomes a slightly optimistic lie.
A worked comparison
Take a single-channel patch with values 100, 150, 200, 250. Mean is 175, standard deviation is about 55.9.
| Raw | Min-max (÷255) | Standardised |
|---|---|---|
| 100 | 0.392 | -1.342 |
| 150 | 0.588 | -0.447 |
| 200 | 0.784 | 0.447 |
| 250 | 0.980 | 1.342 |
Both keep the relative spacing intact. Standardisation centres on zero; min-max keeps everything positive and bounded. Neither is universally better, but mismatching one against a pretrained model is a real error — feeding [0,1] inputs to a network expecting standardised ones typically costs several points of accuracy for no visible reason.
Whitening goes further: it decorrelates the channels so the covariance matrix becomes the identity. It helped older architectures and is largely obsolete now, because batch normalisation layers inside the network achieve a similar effect adaptively, at every layer rather than only the input.
Geometric transformations
These move pixels around the grid. All of them, from a simple shift to a full perspective warp, are matrix multiplications on coordinates.
Rotation
Rotating by angle θ about the origin maps a point as:
Two things bite people here. First, rotation is about the origin, which is the top-left corner of an image — so an unmodified rotation swings the picture out of frame. You need to translate the centre to the origin, rotate, then translate back. cv2.getRotationMatrix2D does that for you if you give it the centre point.
Second, the rotated rectangle does not fit inside the original rectangle. Corners get cut off, and empty triangles appear at the edges. You either accept the crop, expand the canvas, or fill the gaps — and how you fill them matters, because a black triangle is a strong artificial edge that convolution filters respond to. Reflecting the border content usually looks more natural than a constant fill.
Translation, flipping, shearing
Translation shifts by (tx,ty) and leaves an empty strip on one side, with the same fill question.
Flipping mirrors the array — it is pure indexing, costs nothing, and loses no information. But it is not always semantically safe. Horizontal flips are fine for cats and cars. They are wrong for text, for handwritten digits (a mirrored 2 is not a 2), and for medical images where left and right sides mean different organs. Vertical flips are fine for satellite imagery and almost always wrong for street photographs, because upside-down cars do not occur.
Shearing slants the image, as if the top were pushed sideways relative to the bottom. It mimics the mild distortion of viewing a flat object at an angle.
Affine versus perspective
An affine transform combines rotation, scale, translation and shear in a single 2×3 matrix. Its defining property: parallel lines stay parallel. Three corresponding point pairs determine it uniquely.
A perspective transform uses a 3×3 matrix and does not preserve parallelism — which is exactly what you need to model a camera looking at a plane from an angle. Railway tracks converge; an affine transform cannot express that. Four point pairs determine it. This is the operation behind every "scan a document with your phone" feature: find the four corners of the page in the photo, map them to the four corners of a rectangle, and the skewed photograph becomes a flat scan.
1import cv22import numpy as np34# Four corners of a page as they appear in the photo (x, y)5src = np.float32([[142, 91], [510, 130], [548, 640], [98, 590]])6# Where they should end up: a clean A4-ish rectangle7dst = np.float32([[0, 0], [420, 0], [420, 594], [0, 594]])89M = cv2.getPerspectiveTransform(src, dst)10flat = cv2.warpPerspective(photo, M, (420, 594))Filtering: changing pixel values using their neighbours
Geometric transforms move pixels. Filters change them, by replacing each pixel with a function of the pixels around it.
Blurring
A blur replaces each pixel with a weighted average of its neighbourhood. A 3×3 box blur weights all nine equally; a Gaussian blur weights the centre most heavily, falling off with distance, which looks more natural and avoids the boxy artefacts of a plain average.
Blurring sounds destructive, and it is — that is the point. It suppresses high-frequency content, which means sensor noise and irrelevant fine texture. Denoising before edge detection is standard practice, because an edge detector responds to sharp changes and noise is nothing but sharp changes.
Median blur deserves special mention. Instead of averaging, it takes the median of the neighbourhood. Against salt-and-pepper noise — isolated pure-black and pure-white pixels — averaging smears each bad pixel across its neighbours, while the median discards it outright because an extreme value is never the middle value. For that specific noise type the difference is dramatic.
Bilateral filtering smooths within regions but refuses to average across strong edges, by weighting neighbours on both spatial distance and colour similarity. It removes noise while keeping boundaries crisp, at considerably more compute.
Sharpening and edge detection
Sharpening amplifies local differences. A standard kernel:
0 -1 0-1 5 -1 0 -1 0The weights sum to 1, so overall brightness is preserved, while the negative surround subtracts the local average and exaggerates any difference between a pixel and its neighbours. Apply it to a noisy image and you sharpen the noise too — which is why blurring first, then sharpening, is not as contradictory as it sounds.
Edge detection asks where intensity changes fast. Sobel filters estimate the horizontal and vertical gradients; their magnitude Gx2+Gy2 is edge strength. Canny wraps that in a full procedure: Gaussian blur, Sobel gradients, thinning of thick responses to single-pixel lines, then a two-threshold rule where strong edges are kept outright and weak ones survive only if connected to a strong one. That last step is why Canny produces clean connected contours instead of speckle.
Histogram equalisation
A low-contrast photograph has all its pixel values bunched into a narrow band — say 90 to 150, using a quarter of the available range. Equalisation redistributes values so they spread across the full 0–255, using the cumulative distribution of the existing values as the mapping.
Applied globally it often overdoes it, amplifying noise in flat regions such as sky. CLAHE (contrast-limited adaptive histogram equalisation) fixes this by equalising within small tiles, so a dark corner is brightened without washing out an already-bright centre, and by clipping the histogram to cap how much any single region can be amplified.
One important detail: apply equalisation to a lightness channel, not to R, G and B separately. Equalising the three colour channels independently changes their relative proportions and shifts the hue — faces turn orange or green. Convert to LAB, equalise L, convert back.
Order of operations
The same steps in a different order give different results, and some orderings are simply wrong.
| Step | Position | Why there |
|---|---|---|
| Load and fix channel order | 1 | Everything downstream assumes a known layout |
| Denoise | 2 | Before anything that amplifies detail |
| Geometric transforms | 3 | On full resolution, so interpolation has real data to work with |
| Resize / crop | 4 | After geometry, before per-pixel scaling |
| Convert to float | 5 | Avoids integer overflow in the next step |
| Normalise or standardise | 6 | Always last |
Normalisation must come last, and the reason is concrete. Suppose you standardise first, producing values around [−2,2], then resize with a function expecting uint8. It clips your negatives to zero, and half the information is gone. Or you normalise to [0,1] then apply a brightness adjustment written for 0–255 — adding 30 to a value of 0.4 gives 30.4, which is not a brightness adjustment, it is a catastrophe.
Geometry before resizing matters too: rotating a 224×224 image throws away detail that rotating the 4000×3000 original would have preserved, because interpolation on a small grid has almost nothing to interpolate from.
What this means when you build a pipeline
Write the preprocessing as one function, and call the same function at training time and at inference time. The single most common cause of a model that scores well and performs badly is a training pipeline and a serving pipeline that were written separately and drifted apart — a different resize interpolation, a forgotten channel swap, ImageNet statistics on one side and division by 255 on the other. None of these throw an error. They just cost accuracy.
Then look at the output before you trust it. Not the shape — the actual picture. Save a grid of a dozen preprocessed images, denormalise them back to viewable range, and look. Upside-down images, blue faces, masks with impossible class values and objects cropped out of frame are all instantly visible to a human eye and completely invisible in a loss curve.
Finally, keep every decision recorded next to the model weights: input size, interpolation method, channel order, and the exact mean and standard deviation used. Six months from now, someone deploying that model needs to reproduce the numbers exactly. If the record does not exist, they will guess, and their guess will be wrong in a way that costs a few percent and takes a week to find.