- MantraMindAI
- Blog
- Multimodal AI
Computer vision for beginners: from pixel grids to bounding boxes
Jai Rao
August 22, 202618 min read
A ground-up walkthrough of computer vision: what an image is to a computer, why convolution beats dense layers, and how classification, detection and segmentation are scored.
Load a photo in Python and print its shape. You get back something like (1080, 1920, 3) — three small numbers describing six million individual measurements of light. Everything in computer vision starts there. A model never sees a cat, a cracked weld, or a licence plate. It sees a large array of integers and learns which arrangements of those integers tend to arrive with which labels.
Taking that literally explains most of the field's design decisions. The sheer size of the array is why you cannot hand raw pixels to a dense layer. The fact that a cat shifted twenty pixels right is a completely different array is why convolution exists. And the fact that "what is in this image" and "where is it" are different questions is why object detection needs its own labels, its own loss, and its own metric. What follows builds that path from the array upward, with arithmetic you can check on paper and code you can run.
An image is a grid of numbers
A grayscale image is a two-dimensional array: height rows by width columns, one number per pixel giving brightness. Almost always that number is an unsigned 8-bit integer, so it ranges from 0 (black) to 255 (white). A colour image adds a third axis, the channels, usually red, green, and blue. So a colour image is height x width x 3, and the pixel at row 40, column 12 is a triple like [122, 118, 110] — a slightly warm dark grey.
You can confirm all of this in four lines.
import numpy as npfrom PIL import Imageimg = np.asarray(Image.open("photo.jpg"))print(img.shape, img.dtype) # (1080, 1920, 3) uint8print(img.size) # 6220800 individual numbersprint(img[40, 12]) # [122 118 110] -> R, G, BThat last line is the whole point: 6,220,800 numbers, and the model has to find structure in them. Even a modest 224x224 thumbnail, the size most pretrained image models expect, is 224 x 224 x 3 = 150,528 values. Before any of it reaches a network you normally divide by 255 to get floats in the range 0 to 1, then subtract a per-channel mean and divide by a per-channel standard deviation, which keeps activations in a range the optimiser handles well. One more detail trips people up early: NumPy and PIL store images as height x width x channels, while PyTorch expects channels x height x width, with a batch dimension in front. Video adds a time axis on top, which is why it gets expensive so quickly.
Why a fully connected layer falls apart on pixels
The obvious first idea is to flatten the image into one long vector and feed it to an ordinary neural network. Do the arithmetic and the idea dies. Flattening a 224x224 colour image gives 150,528 inputs. Connect those to a hidden layer of just 1,000 units and you have 150,528 x 1,000 = about 150 million weights in that single layer, roughly 600 MB in 32-bit floats. That is one layer, before you have learned anything useful, and every one of those weights needs gradient updates and enough data to constrain it.
Parameter count is only the cheaper of the two problems. The deeper one is that a flattened representation ties every weight to an absolute position. The weight that looks at input index 5,000 sees one specific pixel. Photograph the same object twenty pixels to the left and a completely different set of weights is now responsible for recognising it, and those weights have to learn the pattern all over again from their own examples. Nothing in the architecture says "a vertical edge here means the same thing as a vertical edge there". Flattening also destroys locality: pixel (0, 223) and pixel (1, 0) are neighbouring entries in the flat vector but sit at opposite ends of the image.
Convolution attacks both problems with a single idea — apply the same small set of weights at every position.
Convolution, worked out by hand
A convolutional filter is a small array of weights, typically 3x3, that slides across the image. At each position you multiply the filter elementwise against the patch of pixels underneath it, sum the products, and write that single number to the output. Slide by one pixel, repeat. The output is called a feature map, and it says how strongly the pattern the filter encodes appears at each location.
Here is a 5x5 patch containing a vertical edge — dark on the left, bright on the right — and a filter that detects exactly that.
patch filter 10 10 200 200 200 -1 0 1 10 10 200 200 200 -1 0 1 10 10 200 200 200 -1 0 1 10 10 200 200 200 10 10 200 200 200Take the top-left position. The filter covers rows 0-2 and columns 0-2, so each of its three rows sees the values 10, 10, 200. Multiply and sum: (-1 x 10) + (0 x 10) + (1 x 200) = 190. Three identical rows, so the output is 570. Slide one column right and the window covers 10, 200, 200 per row: (-1 x 10) + (0 x 200) + (1 x 200) = 190 again, so 570 again. Slide once more and the window sees 200, 200, 200 per row, giving (-200) + 0 + 200 = 0. The filter fires where brightness increases left to right and stays silent where the region is flat.
The whole 3x3 output is three identical rows of 570 570 0. Note that the edge sits between two columns of the input but produces two columns of response — a 3-wide filter smears its answer across roughly its own width, which is why deep networks with small filters still see blurry, overlapping evidence at every layer.
This is the entire operation, in runnable form.
import numpy as nppatch = np.array([[10, 10, 200, 200, 200]] * 5, dtype=float)vert = np.array([[-1., 0., 1.]] * 3)def conv2d(img, k): kh, kw = k.shape out = np.zeros((img.shape[0] - kh + 1, img.shape[1] - kw + 1)) for y in range(out.shape[0]): for x in range(out.shape[1]): out[y, x] = np.sum(img[y:y + kh, x:x + kw] * k) return outprint(conv2d(patch, vert)) # [[570 570 0] x 3]print(conv2d(patch, vert.T)) # all zerosThe second print is the interesting one. Transposing the filter turns it into a horizontal edge detector, and applied to a purely vertical edge it returns zeros everywhere, because each column of the patch is constant so -v + 0 + v cancels. Filters are selective — that selectivity is what makes a stack of them informative.
Three consequences of sliding one filter instead of wiring every pixel:
- Weight sharing. A 3x3 filter on a 3-channel input has 3 x 3 x 3 = 27 weights plus one bias, regardless of whether the image is 224 pixels wide or 4,000. Sixty-four such filters cost 1,728 weights and 64 biases. Compare that with 150 million.
- Translation equivariance. Move the edge in the input and the response moves with it, by construction. The pattern is learned once and applies everywhere.
- Local receptive fields. Each output value depends on a small contiguous neighbourhood, so spatial structure is preserved instead of being scrambled by flattening.
Two knobs control the output size. Stride is how far the filter jumps between positions; stride 2 halves the width and height. Padding adds a border of zeros so the output can keep the input's size. For a square input the output side is (input - filter + 2 x padding) / stride + 1. Critically, real networks do not use hand-designed edge filters; the 27 weights are initialised randomly and learned by gradient descent. Edge-like filters in the first layer are something the training process tends to arrive at, not something anyone programs.
Depth, receptive fields, and pooling
A single 3x3 layer can only ever look at 3x3 pixels at once, which is enough for an edge and nowhere near enough for a face. Depth fixes that. Stack a second 3x3 layer on the first and each of its outputs depends on a 3x3 window of first-layer outputs, each of which depended on its own 3x3 window of pixels — so the second layer effectively sees 5x5 of the original image. A third sees 7x7. That growing window is the receptive field, and it is why very deep stacks of tiny filters work: cheap layers compose into wide views.
Growing the receptive field two pixels per layer is slow, so networks also reduce resolution as they go, either by pooling or by strided convolutions. Max pooling with a 2x2 window and stride 2 keeps only the largest value in each 2x2 block, halving height and width and discarding three quarters of the positions. That sounds destructive, and it is, deliberately: what survives is the strongest evidence in each neighbourhood, and the exact pixel where it occurred stops mattering. Halving the resolution also doubles the receptive field of everything above it in the stack, for free.
The standard shape of a classifier follows from this. Spatial size shrinks while channel count grows — 224x224x3 might become 112x112x64, then 56x56x128, on down to something like 7x7x512. Early channels respond to oriented edges and colour blobs; deeper channels respond to textures, then parts, then whole-object configurations, because they are built from combinations of the layers below. At the end you collapse the remaining 7x7 grid with global average pooling, giving one number per channel — a 512-length vector summarising the image — and feed that to a small linear layer that outputs one score per class.
Classification, detection, segmentation: three jobs, three label formats
"Computer vision" covers several tasks that look adjacent and are not interchangeable. They differ in what the model outputs, what your annotations must contain, and how you score results. Getting this choice wrong is the most expensive early mistake, because the labels are the part you cannot cheaply redo.
| Task | Label per image | Model output | Primary metric |
|---|---|---|---|
| Classification | One class name | Score per class | Accuracy, per-class precision/recall |
| Object detection | List of boxes, each with a class | Boxes + classes + confidences | mAP at one or more IoU thresholds |
| Semantic segmentation | A class id for every pixel | Per-pixel class map | Mean IoU, Dice |
| Instance segmentation | A separate mask per object | Masks + classes | Mask mAP |
Classification answers "what is this a picture of" with one label. It is the cheapest to annotate — a click per image — and the cheapest to evaluate, but plain accuracy is a trap whenever classes are imbalanced. If 98 out of 100 inspected parts are fine, a model that always predicts "fine" scores 98% accuracy and finds zero defects. Look at per-class recall and a confusion matrix instead, and decide up front which mistake hurts more: missing a defect, or stopping the line for a good part.
Detection answers "what, and where" by returning a list of axis-aligned boxes, each with a class and a confidence. The output length varies per image, which is why detection needs specialised architectures and losses rather than a single softmax. Annotation cost jumps sharply — someone must drag a box around every instance, and disagreements about how tightly to crop become measurable noise in your metric.
Segmentation goes further and assigns a class to every pixel, which is what you want for area measurements: how much of this crop is diseased, what fraction of this weld is porous. Semantic segmentation labels pixels by class, so two touching cars merge into one "car" region; instance segmentation keeps them separate. Pixel annotation is by far the most laborious of the four.
The practical rule: pick the weakest task that still answers your business question. If the decision is "flag this image for review", classification is enough, and you will get there with a fraction of the labelling effort. Only move up when a box or a mask is genuinely required by what happens next.
IoU, and why accuracy cannot score a box
Accuracy needs a right answer to compare against. A predicted box is almost never pixel-identical to the ground truth, so "correct" has to be defined by overlap. Intersection over Union does that: divide the area where the two boxes overlap by the area they cover together. Identical boxes give 1.0. Boxes that do not touch give 0.0.
Take a ground-truth box spanning x from 10 to 110 and y from 10 to 110, and a prediction offset by twenty pixels in both directions, from 30 to 130. Each box is 100 x 100 = 10,000 square pixels. They overlap over x from 30 to 110 and y from 30 to 110, which is 80 x 80 = 6,400. The union is 10,000 + 10,000 - 6,400 = 13,600, so IoU is 6,400 / 13,600 = 0.47. A prediction that looks nearly right to the eye lands under the conventional 0.5 threshold and is scored as a miss — plus, because the ground-truth object went unmatched, it also counts as a missed detection. IoU is unforgiving on small objects, which is exactly why detection benchmarks report it explicitly.
Nine lines of arithmetic implement it, and the clamping is the part to look at.
def iou(a, b): ax1, ay1, ax2, ay2 = a bx1, by1, bx2, by2 = b ix = max(0, min(ax2, bx2) - max(ax1, bx1)) iy = max(0, min(ay2, by2) - max(ay1, by1)) inter = ix * iy union = (ax2 - ax1) * (ay2 - ay1) + (bx2 - bx1) * (by2 - by1) - inter return inter / unionprint(iou((10, 10, 110, 110), (30, 30, 130, 130))) # 0.470588...print(iou((10, 10, 110, 110), (10, 10, 110, 110))) # 1.0The max(0, ...) calls matter: without them, two disjoint boxes produce negative side lengths that multiply into a bogus positive intersection.
Average precision builds on IoU. Fix a threshold, say 0.5, then walk the predictions from highest confidence to lowest, marking each as a true positive if it matches an unclaimed ground-truth box above that threshold and a false positive otherwise. That traces a precision-recall curve, and AP is the area under it. Mean average precision averages AP across classes, and the common mAP@[.5:.95] also averages over IoU thresholds from 0.5 to 0.95, rewarding tight boxes rather than approximate ones. Segmentation uses the same overlap idea on pixel sets instead of rectangles. One caution: an mAP figure is meaningless without knowing the dataset, the threshold, and the object size distribution, so never compare numbers across different evaluation sets.
Fine-tuning beats training from scratch
Beginners routinely assume the job is to design and train a network from random weights. It almost never is. Edges, textures, and part-like patterns are close to universal across natural images, so a backbone trained on a large general image collection has already learned them, and those layers transfer to your task. Starting from those weights and adapting them is transfer learning, and with a few thousand labelled images it will beat anything you train from scratch on the same data, usually by a wide margin and in a fraction of the time.
The mechanics are short. Load a pretrained model, throw away its final classification layer, bolt on a new one sized to your classes, freeze everything else, and train only the new head.
import torch, torch.nn as nnfrom torchvision import modelsmodel = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1)for p in model.parameters(): p.requires_grad = Falsemodel.fc = nn.Linear(model.fc.in_features, 4) # 4 classes, trainableopt = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)loss_fn = nn.CrossEntropyLoss()model.train()for images, labels in train_loader: opt.zero_grad() loss = loss_fn(model(images), labels) loss.backward() opt.step()Because the frozen backbone is doing all the feature extraction, this trains in minutes on a single GPU and gives you an honest baseline before you spend anything else. Replacing model.fc after the freeze loop is deliberate — the new layer is created with requires_grad=True, so it is the only thing the optimiser can move.
When the head plateaus, unfreeze the last block or two and continue with a much smaller learning rate for the pretrained weights. Large updates to a good backbone destroy what makes it good.
for p in model.layer4.parameters(): p.requires_grad = Trueopt = torch.optim.AdamW([ {"params": model.layer4.parameters(), "lr": 1e-5}, {"params": model.fc.parameters(), "lr": 1e-4},])Two rules make this reliable. Preprocess your images exactly the way the pretrained model was trained — same input size, same normalisation constants — or its features will be evaluated on inputs it never saw. And keep the backbone in the domain it was trained for: features from everyday photographs transfer well to product photos and much less well to X-rays or synthetic-aperture radar, where a domain-specific pretrained model is worth hunting for.
Choosing a backbone
The convolutional families — ResNet, EfficientNet, ConvNeXt — remain excellent defaults, and a small ResNet is the right first thing to try. Vision transformers are the main alternative. A ViT chops the image into fixed patches, commonly 16x16 pixels, flattens each patch into a vector, and treats the sequence with the same self-attention machinery that language models use, so any patch can attend to any other in the first layer rather than waiting for the receptive field to grow. That global view is powerful, but the architecture has no built-in assumption that nearby pixels are related, so it has to learn that from data and is correspondingly hungrier for it. In practice this rarely changes your workflow: you download pretrained weights either way, and the freeze-then-fine-tune loop above is identical apart from the layer names.
Augmentation, and the split that lies to you
With a few thousand images, the model will happily memorise them. Data augmentation counters that by showing a slightly different version of each image every epoch — a random crop, a horizontal flip, a small brightness or contrast shift. The model is forced to rely on properties that survive those changes. In torchvision that means two transform pipelines, and the difference between them matters more than the contents of either.
from torchvision import transformsnorm = transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])train_tf = transforms.Compose([ transforms.RandomResizedCrop(224, scale=(0.7, 1.0)), transforms.RandomHorizontalFlip(), transforms.ColorJitter(brightness=0.2, contrast=0.2), transforms.ToTensor(), norm,])eval_tf = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), norm,])Note that eval_tf has no randomness at all. Augmentation belongs to training only; applying it at evaluation makes your metric jitter run to run for no reason. Choose transforms that reflect variation you actually expect to see. A fixed overhead inspection camera never delivers upside-down parts, so heavy rotation just wastes capacity — while a phone-camera app should absolutely train on rotations and blur. And never flip horizontally when left-right matters, as it does for text, digits, or anything chiral.
The subtler failure is contamination between splits, and it is the single most common reason a model reports 97% on validation and disappoints in production. If your images came from video, consecutive frames are nearly identical; a random split scatters frames from the same second into both train and validation, so the model is graded on pictures it has effectively already seen. The same thing happens with multiple photos of one product, several scans from one patient, or duplicated and resized copies of the same file.
The fix is to split by group rather than by image — by capture session, serial number, patient, camera, or day — so no source appears on both sides.
import numpy as npfrom sklearn.model_selection import GroupShuffleSplitpaths = np.array(all_image_paths)groups = np.array(capture_ids) # one id per source, e.g. video clipsplitter = GroupShuffleSplit(n_splits=1, test_size=0.2, random_state=0)train_idx, val_idx = next(splitter.split(paths, groups=groups))assert not (set(groups[train_idx]) & set(groups[val_idx]))The assertion is the part worth keeping. It fails loudly the day someone adds images without a group id, which is much better than discovering the leak after deployment. Two related habits: deduplicate by file hash before splitting, and hold back a genuinely untouched test set that you look at once, at the end. A validation set you have tuned twenty decisions against has stopped being an unbiased estimate.
What to build first, and how to know it works
A sensible first project looks like this. Collect a few hundred images per class from the camera and lighting you will actually deploy with. Frame the problem as classification even if you eventually want boxes. Split by capture group, fine-tune a small pretrained backbone with the head-only recipe above, and get a number. That whole loop should take an afternoon, and everything afterwards is measured against it.
Then interrogate the result rather than celebrating it. Compare against the trivial baseline of always predicting the majority class; if your model does not clearly beat that, the metric is telling you nothing. Print a confusion matrix and read it — confusion concentrated in one pair of classes usually means those classes are genuinely ambiguous in your images, or your labelling guidelines disagree with themselves. Then open the fifty worst-scored validation images by hand. This is the highest-value hour in the project. You will typically find mislabelled examples, a class that is really two classes, or a shortcut the model latched onto, like a timestamp overlay or a background that only appears for defective parts.
Signals that you have a real result: a small gap between training and validation performance, not a large one; performance that holds up on images captured on a different day or camera; and errors that a human annotator also finds hard. Signals that you do not: near-perfect accuracy on the first try, a metric that swings several points when you change the random seed, or a model whose confidence is high on inputs that contain nothing relevant at all.
When you do improve things, the ordering matters. More and better-labelled data, and augmentation that matches your deployment conditions, will move the number further than swapping architectures. Fix the labels first, fix the split second, tune the model third. A bigger backbone applied to a leaking dataset just produces a more confident wrong answer.