Computer Vision Fundamentals

Image Preprocessing Techniques


A team trains a classifier to spot manufacturing defects on a production line. Validation accuracy: 97%. They deploy it. In the factory it gets about 60% — barely better than guessing between two classes.

Nothing is wrong with the model. The training images came from a folder of 4000×3000 photographs, resized by a script that squashed everything to 224×224 regardless of shape. The live camera produces 1280×720 frames, which squash to a different aspect ratio entirely. A circular defect that appeared as a circle in training arrives at inference as an ellipse. The model has never seen an ellipse.

That is a preprocessing bug, and preprocessing bugs have a signature: the model looks fine on your data and falls apart on anyone else's. They are also, unlike architecture problems, entirely preventable — provided you understand what each transformation actually does to the numbers rather than treating the pipeline as a magic incantation copied from a tutorial.

Preprocessing is everything that happens between reading a file from disk and handing an array to a model. Its job is to make images consistent in size, consistent in numeric range, and free of variation that has nothing to do with the task.

The gap between validation and the factory floorRead imageResize witha fixedinterpolationFit aspect:crop or padScale to 0-1Standardisewith trainmean and std97 percent to 60 percent came from serving images through a different resize than training used.
Preprocessing is part of the model — any step that differs at serving time is a silent distribution shift.

Resizing, and why it is never free

Neural networks with fully connected layers require a fixed input size, and batching requires every image in a batch to have identical dimensions. So resizing is almost always the first step. The question is how to invent pixels that were never measured.

Interpolation: guessing the values in between

Enlarging a 2×2 image to 4×4 means producing 16 values from 4. Twelve of them do not exist and must be estimated. That estimation is interpolation, and the method you choose is a trade between speed and smoothness.

MethodHow it decides a valueSpeedResultUse for
Nearest neighbourCopies the closest original pixelFastestBlocky, hard edgesSegmentation masks and label maps
BilinearWeighted average of the 4 surrounding pixelsFastSmooth, slightly softGeneral-purpose default
BicubicCubic fit over the 16 surrounding pixelsSlowerSharper than bilinearUpscaling where quality matters
LanczosWindowed sinc over a wider neighbourhoodSlowestSharpest, can ring at edgesFinal-quality photographic resizing
AreaAverages all pixels falling into the output cellFastClean, no aliasingDownscaling

Two of those rows are not preferences but rules, and both have a concrete failure attached.

Use nearest neighbour for masks. A segmentation mask holds class identifiers: 0 background, 1 road, 2 car. Bilinear interpolation averages neighbouring values, so a pixel on the boundary between class 0 and class 2 becomes 1 — a road pixel that exists nowhere in the original annotation. Your labels now contain fabricated classes along every object edge. Nearest neighbour copies rather than blends, so identifiers stay identifiers.

Use area for downscaling. Shrinking 1000×1000 to 100×100 with bilinear interpolation samples roughly one pixel out of every hundred and discards the rest. If the image contains a fine repeating pattern — a striped shirt, a mesh grille, text — the sampling grid beats against the pattern and produces aliasing: moiré fringes and stripes that were never there. Area interpolation averages over the whole 10×10 block that maps to each output pixel, so all the information contributes and the artefacts vanish.

Downscale by averaging, upscale by interpolating, and resize masks by copying. Almost every resizing artefact you will ever see comes from breaking one of those three.

Aspect ratio: squash, crop, or pad

A 1600×900 image needs to become 224×224. Three options, three different costs:

StrategyWhat happensCostSensible when
Stretch to fitBoth axes scaled independentlyGeometry is distorted; circles become ellipsesShape genuinely does not matter (texture, some medical modalities)
Centre cropScale the short side to 224, cut the long sideContent at the edges is thrown awayThe subject is reliably central
Letterbox / padScale so the long side is 224, fill the restWasted pixels; the model must learn to ignore the paddingDetection, where every object must survive and boxes must stay valid

The defect-detection failure at the top of this page was a stretch applied inconsistently between training and deployment. Note that stretching was not the real sin — using different aspect ratios in training and production was. Consistency matters more than which option you pick.

Letterboxing is the standard choice for object detection because cropping can remove an object entirely, and stretching invalidates the aspect ratio priors that anchor-based detectors rely on. The padding is usually a neutral grey (114, 114, 114) rather than black, so the border does not look like a strong dark edge to the first convolution layer.

Python
import cv2import numpy as npdef letterbox(img, size=224, fill=114):    h, w = img.shape[:2]    scale = size / max(h, w)    nh, nw = round(h * scale), round(w * scale)    # shrinking -> AREA, growing -> LINEAR    interp = cv2.INTER_AREA if scale < 1 else cv2.INTER_LINEAR    resized = cv2.resize(img, (nw, nh), interpolation=interp)    canvas = np.full((size, size, 3), fill, dtype=np.uint8)    top, left = (size - nh) // 2, (size - nw) // 2    canvas[top:top + nh, left:left + nw] = resized    return canvas, scale, (left, top)   # keep scale/offset to map boxes back

Returning the scale and offset is not optional housekeeping. Without them you cannot convert a predicted box in 224-space back to a box in the original image, and you will end up with detections drawn in the wrong place.

Getting the numbers into a sensible range

Raw pixel values run 0 to 255. Feeding those directly to a network trains badly, and it is worth understanding exactly why rather than accepting "you should normalise" as folklore.

The first layer computes a weighted sum of its inputs. With inputs around 150 and a few hundred of them, that sum is in the tens of thousands before any weight is learned. Push that through a sigmoid or tanh and you land deep in the flat region where the gradient is effectively zero — the layer stops learning. Even with ReLU, which does not saturate, the gradients with respect to the weights are proportional to the input magnitude, so large inputs give large gradients, and the learning rate that keeps the first layer stable is far too small for the rest of the network.

There is a second, subtler problem. If one input dimension typically sits near 200 and another near 10, the loss surface becomes a long narrow valley, and gradient descent zig-zags across it instead of running down it. Equalising the scales makes the valley round and the descent direct.

Large inputs do not merely slow training down. They force a learning rate small enough to keep the first layer stable, and that rate is far too small for every layer after it.

Min-max scaling

x′=x−xmin⁡xmax⁡−xmin⁡x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}}

For 8-bit images this is just division by 255, mapping everything to [0,1][0, 1]. Simple, bounded, and preserves the shape of the distribution exactly. Its weakness is sensitivity to outliers: if a single hot pixel reads 255 in an otherwise dim 16-bit scan, that one pixel sets the maximum and compresses everything real into the bottom of the range.

Standardisation (z-score)

x′=x−μσx' = \frac{x - \mu}{\sigma}

Subtract the mean, divide by the standard deviation. The result has mean 0 and standard deviation 1, and is not bounded — values typically land in roughly [−3,3][-3, 3].

Centring on zero matters more than it looks. If every input to a neuron is positive, every weight's gradient shares the same sign on a given step, so all the weights of that neuron must move up together or down together. Progress happens in zig-zags. Zero-centred inputs let individual weights move in whichever direction they need to.

The critical practical point is which mean and standard deviation you use. For ImageNet-pretrained models, use the ImageNet statistics — the network's learned filters expect inputs distributed that way:

Python
IMAGENET_MEAN = [0.485, 0.456, 0.406]IMAGENET_STD  = [0.229, 0.224, 0.225]

For a model trained from scratch on your own data, compute the statistics from your training set only. Computing them over train plus validation leaks information about the validation set into training, and your validation score becomes a slightly optimistic lie.

A worked comparison

Take a single-channel patch with values 100, 150, 200, 250. Mean is 175, standard deviation is about 55.9.

RawMin-max (÷255)Standardised
1000.392-1.342
1500.588-0.447
2000.7840.447
2500.9801.342

Both keep the relative spacing intact. Standardisation centres on zero; min-max keeps everything positive and bounded. Neither is universally better, but mismatching one against a pretrained model is a real error — feeding [0,1][0, 1] inputs to a network expecting standardised ones typically costs several points of accuracy for no visible reason.

Whitening goes further: it decorrelates the channels so the covariance matrix becomes the identity. It helped older architectures and is largely obsolete now, because batch normalisation layers inside the network achieve a similar effect adaptively, at every layer rather than only the input.

Geometric transformations

These move pixels around the grid. All of them, from a simple shift to a full perspective warp, are matrix multiplications on coordinates.

Rotation

Rotating by angle θ\theta about the origin maps a point as:

[x′y′]=[cos⁡θ−sin⁡θsin⁡θcos⁡θ][xy]\begin{bmatrix} x' \\ y' \end{bmatrix} = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix}

Two things bite people here. First, rotation is about the origin, which is the top-left corner of an image — so an unmodified rotation swings the picture out of frame. You need to translate the centre to the origin, rotate, then translate back. cv2.getRotationMatrix2D does that for you if you give it the centre point.

Second, the rotated rectangle does not fit inside the original rectangle. Corners get cut off, and empty triangles appear at the edges. You either accept the crop, expand the canvas, or fill the gaps — and how you fill them matters, because a black triangle is a strong artificial edge that convolution filters respond to. Reflecting the border content usually looks more natural than a constant fill.

Translation, flipping, shearing

Translation shifts by (tx,ty)(t_x, t_y) and leaves an empty strip on one side, with the same fill question.

Flipping mirrors the array — it is pure indexing, costs nothing, and loses no information. But it is not always semantically safe. Horizontal flips are fine for cats and cars. They are wrong for text, for handwritten digits (a mirrored 2 is not a 2), and for medical images where left and right sides mean different organs. Vertical flips are fine for satellite imagery and almost always wrong for street photographs, because upside-down cars do not occur.

Shearing slants the image, as if the top were pushed sideways relative to the bottom. It mimics the mild distortion of viewing a flat object at an angle.

Affine versus perspective

An affine transform combines rotation, scale, translation and shear in a single 2×3 matrix. Its defining property: parallel lines stay parallel. Three corresponding point pairs determine it uniquely.

A perspective transform uses a 3×3 matrix and does not preserve parallelism — which is exactly what you need to model a camera looking at a plane from an angle. Railway tracks converge; an affine transform cannot express that. Four point pairs determine it. This is the operation behind every "scan a document with your phone" feature: find the four corners of the page in the photo, map them to the four corners of a rectangle, and the skewed photograph becomes a flat scan.

Python
import cv2import numpy as np# Four corners of a page as they appear in the photo (x, y)src = np.float32([[142, 91], [510, 130], [548, 640], [98, 590]])# Where they should end up: a clean A4-ish rectangledst = np.float32([[0, 0], [420, 0], [420, 594], [0, 594]])M = cv2.getPerspectiveTransform(src, dst)flat = cv2.warpPerspective(photo, M, (420, 594))

Filtering: changing pixel values using their neighbours

Geometric transforms move pixels. Filters change them, by replacing each pixel with a function of the pixels around it.

Blurring

A blur replaces each pixel with a weighted average of its neighbourhood. A 3×3 box blur weights all nine equally; a Gaussian blur weights the centre most heavily, falling off with distance, which looks more natural and avoids the boxy artefacts of a plain average.

Blurring sounds destructive, and it is — that is the point. It suppresses high-frequency content, which means sensor noise and irrelevant fine texture. Denoising before edge detection is standard practice, because an edge detector responds to sharp changes and noise is nothing but sharp changes.

Median blur deserves special mention. Instead of averaging, it takes the median of the neighbourhood. Against salt-and-pepper noise — isolated pure-black and pure-white pixels — averaging smears each bad pixel across its neighbours, while the median discards it outright because an extreme value is never the middle value. For that specific noise type the difference is dramatic.

Bilateral filtering smooths within regions but refuses to average across strong edges, by weighting neighbours on both spatial distance and colour similarity. It removes noise while keeping boundaries crisp, at considerably more compute.

Sharpening and edge detection

Sharpening amplifies local differences. A standard kernel:

Text
 0  -1   0-1   5  -1 0  -1   0

The weights sum to 1, so overall brightness is preserved, while the negative surround subtracts the local average and exaggerates any difference between a pixel and its neighbours. Apply it to a noisy image and you sharpen the noise too — which is why blurring first, then sharpening, is not as contradictory as it sounds.

Edge detection asks where intensity changes fast. Sobel filters estimate the horizontal and vertical gradients; their magnitude Gx2+Gy2\sqrt{G_x^2 + G_y^2} is edge strength. Canny wraps that in a full procedure: Gaussian blur, Sobel gradients, thinning of thick responses to single-pixel lines, then a two-threshold rule where strong edges are kept outright and weak ones survive only if connected to a strong one. That last step is why Canny produces clean connected contours instead of speckle.

Histogram equalisation

A low-contrast photograph has all its pixel values bunched into a narrow band — say 90 to 150, using a quarter of the available range. Equalisation redistributes values so they spread across the full 0–255, using the cumulative distribution of the existing values as the mapping.

Applied globally it often overdoes it, amplifying noise in flat regions such as sky. CLAHE (contrast-limited adaptive histogram equalisation) fixes this by equalising within small tiles, so a dark corner is brightened without washing out an already-bright centre, and by clipping the histogram to cap how much any single region can be amplified.

One important detail: apply equalisation to a lightness channel, not to R, G and B separately. Equalising the three colour channels independently changes their relative proportions and shifts the hue — faces turn orange or green. Convert to LAB, equalise L, convert back.

Order of operations

The same steps in a different order give different results, and some orderings are simply wrong.

StepPositionWhy there
Load and fix channel order1Everything downstream assumes a known layout
Denoise2Before anything that amplifies detail
Geometric transforms3On full resolution, so interpolation has real data to work with
Resize / crop4After geometry, before per-pixel scaling
Convert to float5Avoids integer overflow in the next step
Normalise or standardise6Always last

Normalisation must come last, and the reason is concrete. Suppose you standardise first, producing values around [−2,2][-2, 2], then resize with a function expecting uint8. It clips your negatives to zero, and half the information is gone. Or you normalise to [0,1][0, 1] then apply a brightness adjustment written for 0–255 — adding 30 to a value of 0.4 gives 30.4, which is not a brightness adjustment, it is a catastrophe.

Geometry before resizing matters too: rotating a 224×224 image throws away detail that rotating the 4000×3000 original would have preserved, because interpolation on a small grid has almost nothing to interpolate from.

What this means when you build a pipeline

Write the preprocessing as one function, and call the same function at training time and at inference time. The single most common cause of a model that scores well and performs badly is a training pipeline and a serving pipeline that were written separately and drifted apart — a different resize interpolation, a forgotten channel swap, ImageNet statistics on one side and division by 255 on the other. None of these throw an error. They just cost accuracy.

Then look at the output before you trust it. Not the shape — the actual picture. Save a grid of a dozen preprocessed images, denormalise them back to viewable range, and look. Upside-down images, blue faces, masks with impossible class values and objects cropped out of frame are all instantly visible to a human eye and completely invisible in a loss curve.

Finally, keep every decision recorded next to the model weights: input size, interpolation method, channel order, and the exact mean and standard deviation used. Six months from now, someone deploying that model needs to reproduce the numbers exactly. If the record does not exist, they will guess, and their guess will be wrong in a way that costs a few percent and takes a week to find.