Course Content
Computer Vision Fundamentals
3 sections · 8 lessons
Pixels, Channels, and Color Spaces
Here is a task that sounds trivial. You have a folder of photographs of traffic scenes, and you want to find the ones containing a red stop sign. Red is red. How hard can it be?
You write the obvious thing: load the image, look at every pixel, and keep the ones where the red value is high and the green and blue values are low. You run it on a bright midday photo and it works beautifully. Then you run it on the rest of the folder.
A photo taken at dusk finds nothing — the sign is there, but the whole scene is dim, so "red is high" is false for every pixel on it. A photo taken in direct sun finds nothing either, because the sign is blown out towards white and now green and blue are high too. A photo of a brick wall lights up like a Christmas tree. And one photo comes back with the sky flagged as red, which makes no sense at all until you discover that the library you used loads channels in the order blue-green-red, not red-green-blue, so you have been testing the blue channel this whole time.
Every one of those failures comes from the same root cause: you do not yet know what an image actually is to a computer. Not metaphorically — literally, as numbers in memory. Get that right and all four bugs become obvious. Get it wrong and you will keep writing code that works on one photo and fails on a hundred.
An image is a grid of numbers, and nothing else
Strip away the file format, the compression, the viewer application. What remains is a rectangular grid of measurements. Each cell in the grid is a pixel — short for "picture element" — and it holds a number, or a small group of numbers, describing how much light was measured at that point on the sensor.
For a black-and-white image, one number per pixel is enough. That number is intensity: 0 means no light at all (black), and the maximum value means as much light as the sensor can record (white). Everything between is a shade of grey.
Here is a genuine 5×5 image, written out in full:
0 0 64 0 0 0 64 128 64 0 64 128 255 128 64 0 64 128 64 0 0 0 64 0 0That is a small bright cross on a dark background. There is no image hiding behind those numbers — the numbers are the image. When a viewer shows it to you, it maps each number to a shade and paints a square. That is the whole trick.
Every operation in computer vision — every filter, every neural network, every detector — is arithmetic on grids of numbers like this one. If you can hold that in your head, nothing that follows is mysterious.
The indexing trap: rows before columns
This is the single most common source of silent bugs for people new to the field, and it is worth a full minute of attention.
In school mathematics you learned to write a point as (x,y): horizontal first, vertical second. Graphics libraries and drawing APIs mostly follow that convention. But an image in memory is a two-dimensional array, and arrays are indexed row first. A row is a horizontal slice, so the row index is the vertical position.
| Convention | First index | Second index | Origin | Typical users |
|---|---|---|---|---|
| Array indexing | row = y (down) | column = x (right) | top-left | NumPy, PyTorch, OpenCV arrays |
| Cartesian / drawing | x (right) | y (down) | top-left | cv2.circle, PIL drawing, most GUI code |
| Maths class | x (right) | y (up) | bottom-left | textbooks, plotting axes |
So img[10, 50] is the pixel 10 rows down and 50 columns across. But cv2.circle(img, (10, 50), ...) draws at 10 across and 50 down. The same pair of numbers means two different places depending on which API you hand it to.
What goes wrong if you ignore this: on a square image, nothing visible — the bug hides completely. On a non-square image you get an index-out-of-range error if you are lucky, and a silently transposed result if you are not. The fix is a habit, not a rule: whenever you write a coordinate, say out loud which convention the function on the receiving end expects.
Shape: how the grid is described
The shape of an image array is the tuple of its dimensions. A 1920×1080 colour photograph has shape (1080, 1920, 3) — height, then width, then channels. Note that the human phrase "1920 by 1080" puts width first while the array puts height first. Another place to be careful.
Deep learning frameworks add a fourth dimension for the batch, and then disagree with each other about the order:
| Layout | Shape | Used by | Why |
|---|---|---|---|
| NHWC ("channels last") | (32, 224, 224, 3) | TensorFlow, most image libraries | Matches how images are stored on disk; the three channel values of one pixel sit next to each other in memory |
| NCHW ("channels first") | (32, 3, 224, 224) | PyTorch | Faster for convolution on GPUs, because a whole channel is contiguous |
If you feed a PyTorch model an array shaped (224, 224, 3), it will interpret 224 as the number of channels and complain, or worse, quietly do arithmetic that means nothing. tensor.permute(2, 0, 1) is the conversion you will type a thousand times.
Channels: stacking grids to make colour
One grid of numbers gives you grey. Colour needs three, because human colour vision has three types of cone cell, responding roughly to long, medium and short wavelengths. A colour image is therefore three grids stacked on top of each other, called channels. Red, green and blue.
Each pixel now holds three numbers, and the colour you perceive is the mixture. Let us build a 2×2 image by hand and read it back.
Red channel Green channel Blue channel255 0 0 255 0 0255 255 255 0 0 255Reading pixel by pixel:
| Position (row, col) | R | G | B | Colour |
|---|---|---|---|---|
| (0, 0) | 255 | 0 | 0 | Pure red |
| (0, 1) | 0 | 255 | 0 | Pure green |
| (1, 0) | 255 | 255 | 0 | Yellow — red and green together |
| (1, 1) | 255 | 0 | 255 | Magenta |
Yellow catching people out is normal. In paint, mixing red and green gives mud. On a screen you are mixing light, not pigment, and light adds. Red light plus green light lands on your retina as yellow. This is why the model is called additive: start at black (0, 0, 0) and add light to get towards white (255, 255, 255).
Some images carry a fourth channel, alpha, storing opacity rather than colour: 0 is fully transparent, 255 fully opaque. PNG supports it; JPEG does not. A common bug is loading a PNG expecting shape (H, W, 3) and getting (H, W, 4), which then breaks the first matrix multiplication downstream.
Colour spaces: the same colour, different coordinates
RGB is not "what colour is". It is one coordinate system for describing colour, chosen because it matches how screens emit light. There are others, and choosing the right one turns hard problems into easy ones — which is exactly what went wrong in the stop-sign task at the start.
Why RGB fights you
Consider the same red sign photographed three times:
| Lighting | R | G | B |
|---|---|---|---|
| Bright noon sun | 230 | 40 | 45 |
| Overcast afternoon | 150 | 28 | 30 |
| Deep shade at dusk | 70 | 14 | 15 |
All three are the same physical red. But all three numbers changed, and they changed a lot. A rule like R > 200 catches the first and misses the other two. There is no threshold on R, G and B that captures "red" across lighting, because in RGB, brightness is smeared across all three channels. Change the light and every coordinate moves.
HSV: pull brightness out into its own channel
HSV re-describes the same colour with three different quantities:
- Hue — which colour it is, as an angle around a colour wheel. Red sits at 0°, green at 120°, blue at 240°. OpenCV squeezes this into 0–179 so it fits in a byte, so red is near 0 or near 179.
- Saturation — how vivid it is, from 0 (grey) to maximum (pure colour).
- Value — how bright it is, from 0 (black) to maximum.
Convert those three sign pixels to HSV and something useful happens:
| Lighting | H | S | V |
|---|---|---|---|
| Bright noon sun | 179 | 211 | 230 |
| Overcast afternoon | 0 | 207 | 150 |
| Deep shade at dusk | 179 | 204 | 70 |
Hue is essentially constant — 179 and 0 are neighbours, because red sits exactly where OpenCV's hue scale wraps around (which is also why a red mask in OpenCV needs two hue ranges, one near 0 and one near 179). Saturation barely moves. Only Value changes — which is correct, because the only thing that actually changed was how much light there was. Now the rule is easy: hue near red, saturation reasonably high, and ignore value entirely. That single change of coordinates is the difference between a detector that works in one photo and one that works in a hundred.
Choosing a colour space is not decoration. It decides which real-world variations show up as changes in your numbers, and therefore which problems your code has to solve.
Grayscale: throwing colour away on purpose
Collapsing three channels to one cuts data by two-thirds and often costs nothing, because edges, texture and shape live in intensity, not colour. The conversion is a weighted average:
The weights are not equal, and the reason is biological: the human eye is far more sensitive to green than to blue. A plain average (R+G+B)/3 makes blue skies look implausibly bright and green foliage look too dark. Work through pure green, (0,255,0): the weighted formula gives 0.587×255≈150, a light grey, while the plain average gives 85, a dark grey. Ask anyone which better matches how bright green looks and they will pick 150.
Convert to grey when colour genuinely carries no information for your task — document scanning, most edge detection, classic feature matching. Do not convert when colour is the signal: ripe versus unripe fruit, red versus green traffic lights, skin-tone analysis.
LAB and YCbCr
LAB stores lightness (L) plus two colour axes: A runs green to red, B runs blue to yellow. Its selling point is perceptual uniformity: a fixed numeric distance corresponds to roughly the same perceived difference anywhere in the space. In RGB that is badly false — two dark blues 20 units apart look identical, two mid-greens 20 units apart look clearly different. So if you need to measure "how different are these two colours", LAB gives an answer that matches human judgement.
YCbCr splits luma (Y, brightness) from two chroma channels. This is the basis of JPEG and essentially all video compression, and the reason is a fact about your eyes: they resolve fine detail in brightness far better than in colour. So codecs keep Y at full resolution and store Cb and Cr at half or quarter resolution. You lose colour detail you were never going to notice, and the file shrinks dramatically.
| Space | Channels | Reach for it when |
|---|---|---|
| RGB | Red, Green, Blue | Displaying images; feeding a neural network |
| Grayscale | Intensity | Colour is irrelevant and you want speed |
| HSV | Hue, Saturation, Value | Colour-based thresholding under changing light |
| LAB | Lightness, A, B | Measuring colour difference; lightness-only contrast fixes |
| YCbCr | Luma, Cb, Cr | Compression; anything touching video pipelines |
The BGR trap
OpenCV loads images as blue, green, red. It has done so since the late 1990s for reasons of camera hardware conventions at the time, and it will never change because too much code depends on it. Almost everything else — PIL, matplotlib, scikit-image, PyTorch's model zoo — uses RGB.
The symptom is unmistakable once you know it: skin looks blue, sky looks orange, and everyone in the photo appears to have been rescued from a cold lake. The fix is one line, cv2.cvtColor(img, cv2.COLOR_BGR2RGB), but the real damage happens when there is no visible display step. If you load with OpenCV and feed straight into a model trained on RGB, the image looks fine to nobody because nobody looks at it, and accuracy just quietly drops by several points while you hunt for the problem in your training loop.
Bit depth: how finely each number is measured
Bit depth is how many bits store one channel value, which sets how many distinct levels exist.
| Depth | Range | Levels | Where you meet it |
|---|---|---|---|
| 8-bit | 0–255 | 256 | JPEG, PNG, almost every consumer photo |
| 16-bit | 0–65535 | 65,536 | Medical scans, satellite imagery, camera raw |
| 32-bit float | 0.0–1.0 typically | continuous in practice | Intermediate results, HDR, model inputs |
Eight bits is enough for viewing because the eye cannot pick out more than roughly 200 grey levels in one scene. It is not always enough for processing. A CT scan spans thousands of meaningful density values; squash it to 256 and distinctions a radiologist needs disappear.
The practical danger with 8-bit integers is overflow. Brightening an image by adding 50 to every pixel, in NumPy with dtype uint8, turns a pixel at 220 into 270 — which wraps around to 14. Your bright sky develops black patches. The fix is to convert to a wider type, do the arithmetic, clip to the valid range, and convert back:
1import numpy as np23img = np.array([[200, 220, 250]], dtype=np.uint8)45wrong = img + 50 # [250, 14, 44] -- wrapped around6right = np.clip(img.astype(np.int16) + 50, 0, 255).astype(np.uint8)7 # [250, 255, 255] -- saturates correctlyStorage formats and what they cost you
The format on disk is not the image; it is an encoding of it. What matters for vision work is what the encoding destroys.
| Format | Compression | Alpha | Notes for vision work |
|---|---|---|---|
| JPEG | Lossy | No | Small files, but discards high-frequency detail and leaves 8×8 block artefacts. Re-saving repeatedly degrades further. |
| PNG | Lossless | Yes | Pixel-exact. Correct choice for masks, labels and anything you will compute on. |
| TIFF | Either | Yes | Handles 16-bit and multi-page. Standard in microscopy and geospatial work. |
| WebP | Either | Yes | Smaller than JPEG at equal quality; less universally supported by older tooling. |
The rule worth internalising: never store a segmentation mask as JPEG. A mask is a grid of class identifiers — 0 for background, 1 for road, 2 for car. JPEG's lossy compression will shift a 1 to a 0 or a 2 near every boundary, inventing classes that were never in your labels. The file will still open. Your training will just be subtly wrong, and the error is invisible until you plot the histogram of label values and find fractional classes that should not exist.
Seeing it for yourself
Reading about this is far less convincing than printing the numbers. This snippet builds a small image from scratch, inspects it, and demonstrates the lighting problem concretely:
1import numpy as np2import cv234# Build a 2x2 image by hand, in RGB order5rgb = np.array([6 [[255, 0, 0], [ 0, 255, 0]],7 [[255, 255, 0], [255, 0, 255]],8], dtype=np.uint8)910print(rgb.shape) # (2, 2, 3) -> height, width, channels11print(rgb[1, 0]) # [255 255 0] -> row 1, col 0 is yellow1213# The same red under three lighting conditions14reds_rgb = np.array([[[230, 40, 45], [150, 28, 30], [70, 14, 15]]], dtype=np.uint8)1516# cvtColor expects BGR, so reverse the last axis before converting17reds_hsv = cv2.cvtColor(reds_rgb[:, :, ::-1], cv2.COLOR_BGR2HSV)18print(reds_hsv[0])19# Hue stays at the red wrap-around (179 or 0), saturation barely moves; only V drops.2021# Grayscale with perceptual weights vs a naive mean22green = np.array([0, 255, 0], dtype=np.float32)23print(0.299 * green[0] + 0.587 * green[1] + 0.114 * green[2]) # ~149.724print(green.mean()) # 85.0What this changes about how you build
Three habits follow directly from everything above, and they will save you more debugging time than any clever algorithm.
Print the shape and dtype before you trust anything. One line — print(img.shape, img.dtype, img.min(), img.max()) — catches channels-first/channels-last mix-ups, an unexpected alpha channel, an image already scaled to 0–1 when you assumed 0–255, and an all-zero array from a failed file read. Most "the model is not learning" mysteries are one of these five things.
Pick the colour space to match the variation you want to ignore. If lighting changes and you do not care, work in a space that isolates brightness in one channel and drop it. If you are comparing colours numerically, use LAB so the distances mean something. If you are feeding a pretrained network, use RGB, because that is what it was trained on.
Treat the loading step as part of the model. Channel order, bit depth and file format are not incidental plumbing — they are assumptions baked into every downstream computation. A model trained on RGB and served BGR is not a slightly wrong model; it is a model receiving a systematically corrupted input on every single request. Write the loader once, test it by displaying an image and confirming the sky is blue, and then never touch it casually again.