Course Content
Computer Vision Fundamentals
3 sections · 8 lessons
Introduction to Object Detection and Segmentation
You have a classifier that reads a photograph and outputs one of a thousand labels. Point it at a street scene and it says "car" with 94% confidence. Correct — there is a car. There are also four other cars, two pedestrians, a cyclist and a bus.
Now try to use that output for something. A self-driving system needs to know where the pedestrian is and how far away. A retail analytics system needs to count people, not confirm that at least one exists. A medical system needs the outline of the tumour, not the news that a tumour is present.
Classification answers "what is in this image?" with a single label for the whole frame. Almost every real application needs one of three harder questions instead:
| Task | Question answered | Output per image |
|---|---|---|
| Classification | What is this? | One label |
| Object detection | What is here, and where? | A variable-length list of boxes with labels |
| Semantic segmentation | What class is each pixel? | A label map the size of the image |
| Instance segmentation | Which pixels belong to which object? | One mask per object |
The jump from the first row to the second is not a small one, and the reason is worth understanding before any architecture makes sense.
Why detection is structurally harder
A classifier has a fixed output shape: 1000 numbers, one per class, every time. You can train it with a straightforward loss because the target always has the same shape as the prediction.
A detector's output has variable length. This photograph has 3 objects, the next has 47, the next has none. A neural network produces a fixed-size tensor. You cannot simply have an output layer of size "however many objects there happen to be".
Then there is the matching problem. If the model predicts 5 boxes and the ground truth has 3, which prediction should be compared against which label? Get the assignment wrong and the loss punishes a correct prediction for being in the wrong slot.
Every detection architecture is essentially a different answer to those two problems.
Boxes, and the coordinate formats that cause bugs
A bounding box is the smallest axis-aligned rectangle containing an object. Four numbers describe it — but there are three incompatible conventions for which four:
| Format | Numbers | Used by |
|---|---|---|
Corner (xyxy) | x1,y1,x2,y2 — top-left and bottom-right | Pascal VOC, torchvision, most IoU code |
Corner-size (xywh) | x1,y1,w,h — top-left plus dimensions | COCO annotations |
Centre (cxcywh) | cx,cy,w,h — centre plus dimensions, usually normalised to [0,1] | YOLO |
Mixing these is the single most common bug in detection code, and it is nasty because it does not crash. Feed xywh boxes into a function expecting xyxy and it will happily compute a rectangle from (10, 20) to (50, 30) when you meant a 50×30 box at (10, 20). Training proceeds, loss decreases, and the boxes come out systematically wrong. Write conversion functions once, name them explicitly, and never pass a raw array of four numbers between modules without saying which format it is in.
IoU: the measure of "close enough"
A predicted box will never exactly match the ground truth. Intersection over Union quantifies the overlap:
Work an example. Ground truth spans (10,10) to (50,50); prediction spans (20,20) to (60,60). Both are 40×40, so each has area 1600.
- Intersection: x from 20 to 50, y from 20 to 50 → 30×30=900
- Union: 1600+1600−900=2300
- IoU: 900/2300=0.391
1def iou(a, b):2 """Both boxes in xyxy format."""3 x1 = max(a[0], b[0]); y1 = max(a[1], b[1])4 x2 = min(a[2], b[2]); y2 = min(a[3], b[3])56 inter = max(0, x2 - x1) * max(0, y2 - y1) # max(0, ...) handles no overlap7 area_a = (a[2] - a[0]) * (a[3] - a[1])8 area_b = (b[2] - b[0]) * (b[3] - b[1])9 union = area_a + area_b - inter10 return inter / union if union > 0 else 0.0The max(0, ...) clamps are not defensive padding. Without them, two non-overlapping boxes produce negative widths whose product is positive, and you get a confidently wrong non-zero IoU for boxes that do not touch.
By convention IoU ≥ 0.5 counts as a correct detection, though 0.5 is a fairly generous bar — it permits a visibly misplaced box.
mAP: the standard score
Mean average precision is the metric detection papers report, and it is built up in three steps. For one class, sort all predictions by confidence and walk down the list, marking each as correct if it matches an unclaimed ground-truth box above the IoU threshold. This traces a precision-recall curve, and the area under it is the average precision for that class. Average that over all classes and you have mAP.
Notation matters when comparing numbers. mAP@0.5 uses a single IoU threshold of 0.5. mAP@[0.5:0.95], the COCO standard, averages mAP computed at ten thresholds from 0.5 to 0.95 in steps of 0.05, so it rewards precise localisation and produces much lower numbers. A model reporting 55 mAP under COCO and one reporting 80 mAP@0.5 may be the same model.
Non-maximum suppression
Detectors produce many overlapping boxes for the same object — a face might attract twelve predictions at slightly different positions and scales, since several nearby anchors all fired. You want one.
Non-maximum suppression is the standard cleanup, and it is greedy and simple:
- Sort all boxes by confidence, highest first.
- Take the top box and keep it.
- Delete every remaining box whose IoU with it exceeds a threshold, typically 0.45.
- Repeat with the next surviving box until none remain.
1def nms(boxes, scores, thresh=0.45):2 order = sorted(range(len(boxes)), key=lambda i: scores[i], reverse=True)3 keep = []4 while order:5 i = order.pop(0)6 keep.append(i)7 order = [j for j in order if iou(boxes[i], boxes[j]) <= thresh]8 return keepNMS must run per class. Run it across all classes at once and a correctly detected person standing in front of a correctly detected car will have one of the two suppressed, because the boxes overlap heavily. They are different objects; the overlap is real and legitimate.
The threshold is a genuine trade-off with a visible failure at each end. Too low and you delete true detections in crowded scenes — a dense crowd of people produces heavily overlapping boxes that are all correct. Too high and you keep duplicates. Soft-NMS is the middle path: instead of deleting overlapping boxes it reduces their confidence in proportion to the overlap, which lets a strongly-supported box in a crowd survive.
Almost every complaint that "the detector misses people in crowds" is an NMS threshold problem rather than a model problem. Check it before retraining anything.
Two families of detector
Two-stage: propose, then classify
Faster R-CNN splits the job. A region proposal network slides over the feature map and outputs a few thousand candidate regions that might contain something — no class, just "object or not". A second stage then crops the features for each proposal, classifies it, and refines the box.
The variable-length problem is solved by fixing the number of proposals. The matching problem is solved by assigning each proposal to a ground-truth box by IoU during training.
This is accurate, because the second stage examines each candidate carefully with features cropped precisely to it. It is also slow, because that second stage runs hundreds of times per image.
Single-stage: predict everything at once
YOLO takes a different route. Divide the image into a grid; each cell predicts a fixed number of boxes with confidences and class scores, in one forward pass. There is no proposal stage and no per-region processing.
Variable length is handled by predicting a fixed large number of boxes — most with near-zero confidence — and filtering afterwards. Matching is handled by assigning each ground-truth object to the grid cell containing its centre.
SSD sits in the same family but predicts from feature maps at several resolutions, so large objects are found in the coarse deep maps and small ones in the finer shallow maps. This directly addresses the classic single-stage weakness.
| YOLO | Faster R-CNN | SSD | |
|---|---|---|---|
| Stages | One | Two | One |
| Typical speed | 30–150 FPS | 5–15 FPS | 20–60 FPS |
| Accuracy | Good, now near-parity | Historically best | Moderate |
| Small objects | Historically weak | Strong | Moderate |
| Reach for it when | Real time, video, edge devices | Offline accuracy matters most | Balanced embedded use |
The historical accuracy gap has largely closed. Modern single-stage detectors with better loss functions and training recipes match two-stage accuracy at several times the speed, which is why real deployments overwhelmingly use them.
Anchors, and doing without them
Most of these detectors rely on anchors: predefined boxes of assorted sizes and aspect ratios tiled across the image. The network does not predict a box from nothing; it predicts an offset from the nearest anchor, which is a much easier regression problem.
The cost is a set of hyperparameters that must match your data. Anchors tuned for COCO photographs are a poor fit for, say, aerial imagery of ships, which are long thin objects at scales COCO never contains. Symptom: the detector consistently misses a whole category of object shape.
Anchor-free detectors sidestep this. FCOS predicts, for every pixel, the four distances to the object's edges. CenterNet predicts object centres as peaks in a heatmap and regresses size at each peak. DETR treats detection as set prediction with a transformer, matching predictions to targets with a bipartite assignment and removing the need for NMS entirely.
Segmentation: labels for every pixel
A box is a crude description of shape. For a curved road, an irregular tumour or a person's silhouette, you want the actual pixels.
The three flavours
Picture a street scene with three cars and some road.
| Type | What you get | The three cars are… |
|---|---|---|
| Semantic | Class label per pixel | One connected "car" region — indistinguishable |
| Instance | Separate mask per object | Three separate masks (road is not labelled at all) |
| Panoptic | Both combined | Three car instances and a labelled road region |
The distinction between "things" (countable objects: cars, people) and "stuff" (uncountable regions: road, sky, grass) is the reason all three exist. Instance segmentation handles things and ignores stuff. Semantic segmentation handles both but cannot count. Panoptic does everything and is correspondingly harder.
The architecture problem
Semantic segmentation needs a full-resolution output, but the classification backbones that produce good features downsample aggressively — a 224×224 input becomes a 7×7 feature map. You need the semantic richness of the deep layers and the spatial precision of the shallow ones.
Three answers, all still in use:
- Encoder-decoder (U-Net) — downsample as usual, then upsample back to full size, with skip connections carrying the high-resolution early feature maps directly across to the matching decoder stage. Without those skips the boundaries come out hopelessly blurry, because the decoder has to invent detail that was discarded. U-Net was designed for biomedical images and remains the default there.
- Dilated convolutions (DeepLab) — stop downsampling after a point and use dilated kernels instead, which enlarge the receptive field without reducing resolution. Combined with parallel dilation rates, this captures context at several scales at once.
- Feature pyramids — build a pyramid of feature maps at several resolutions and predict from all of them, merging coarse semantics with fine detail.
Instance segmentation
Mask R-CNN is the canonical approach, and it is pleasingly simple in concept: take Faster R-CNN and add a third output branch that predicts a binary mask inside each detected box. Detection gives you the instances; the mask branch gives you the shape within each.
Its one important technical fix is RoIAlign. The earlier RoIPool rounded region coordinates to integer feature-map cells, misaligning the crop by up to half a cell. For classification that hardly matters. For a pixel-accurate mask it produces visibly wrong boundaries. RoIAlign uses bilinear interpolation instead of rounding, and that change alone gave a substantial jump in mask quality.
Metrics
Segmentation is scored with mean IoU: compute IoU per class between predicted and true masks, then average over classes. Averaging over classes rather than pixels matters enormously in imbalanced scenes — road might be 40% of pixels and traffic signs 0.1%. Pixel accuracy would let a model that ignores signs entirely score 95%; mean IoU exposes it immediately.
Average IoU over classes, never over pixels. A model that ignores every rare class outright can still post 95% pixel accuracy on a road scene.
1import torch23def mean_iou(pred, target, num_classes, ignore_index=255):4 """pred, target: integer class maps of the same shape."""5 ious = []6 valid = target != ignore_index7 for c in range(num_classes):8 p, t = (pred == c) & valid, (target == c) & valid9 union = (p | t).sum().item()10 if union == 0:11 continue # class absent in both: skip, do not score 012 ious.append((p & t).sum().item() / union)13 return sum(ious) / len(ious) if ious else 0.0Skipping absent classes rather than scoring them zero is the detail people get wrong. Counting a class that appears in neither prediction nor label as a zero drags the average down for a correct prediction, and makes scores incomparable between images.
Choosing an approach for a real problem
Start from what the downstream system actually consumes, because each step up in output richness costs annotation effort, compute and accuracy.
| What you need to know | Task | Annotation cost per image |
|---|---|---|
| Is this present at all? | Classification | Seconds |
| How many, and roughly where? | Detection | ~1 minute |
| What fraction of the area is X? | Semantic segmentation | ~10–60 minutes |
| The exact shape of each object | Instance segmentation | ~30–90 minutes |
Pixel-level annotation is often one to two orders of magnitude more expensive than boxes. If a box would answer your question, use boxes. Teams routinely commission segmentation labels, spend months on them, and then discover that a counting task only ever needed detection.
Once the task is fixed, start from a pretrained detector rather than a custom architecture. A COCO-pretrained model fine-tuned on a few thousand of your own images will beat anything you design from scratch, and it will beat it in an afternoon rather than a month.
Then tune the parts that are free before touching the model. Confidence threshold and NMS threshold are inference-time settings, cost nothing to change, and often move real-world performance more than a percentage point of mAP would. Missing objects in crowds means raising the NMS threshold. Too many false positives means raising the confidence threshold. Missing small objects usually means the input resolution is too low, not that the model is wrong — a 224×224 input reduces a distant pedestrian to a handful of pixels, and no architecture recovers information that was thrown away in the resize.
Finally, look at the failures rather than the metric. mAP is a single number averaged over classes, thresholds and images, and it hides everything interesting. Plot the failures: which classes, which object sizes, which lighting conditions. Detection problems are almost always concentrated — one class, one scale, one scenario — and the fix is usually more data of that specific kind rather than a better architecture.