Course Content
Computer Vision Fundamentals
3 sections · 8 lessons
Mini Project: CIFAR-10 Image Classification with CNNs
CIFAR-10 is 60,000 colour images, 32×32 pixels each, sorted into ten classes: aeroplane, car, bird, cat, deer, dog, frog, horse, ship, lorry. Fifty thousand for training, ten thousand for testing. It has been the standard proving ground for image classification since 2009, which means there is a well-documented performance ladder to climb and no ambiguity about whether your number is any good.
Look at one of the images at actual size and the difficulty becomes obvious. Thirty-two pixels across is roughly the size of a favicon. A cat at that resolution is a brown smudge with two darker smudges that might be ears. Humans score about 94% on this dataset (Andrej Karpathy's informal 2011 estimate), and they find it genuinely hard.
Here is the ladder you are aiming to climb. The figures are typical ranges from common training recipes, not exact targets — yours will move with the number of epochs and the schedule:
| Approach | Test accuracy | What it tells you |
|---|---|---|
| Random guessing | 10% | The floor |
| Logistic regression on raw pixels | ~40% | How much a linear model can extract |
| Small CNN, no augmentation | ~72% | What convolution alone buys you |
| Same CNN with augmentation and batch norm | ~85% | What training technique buys you |
| ResNet-18 trained from scratch | ~93–95% | What architecture buys you |
| ImageNet-pretrained ResNet, fine-tuned | ~96% | What borrowed features buy you |
| Human | ~94% | Context for all of the above |
The point of this project is not to reach the top of that table. It is to walk up it deliberately, one rung at a time, and be able to explain why each change produced the gain it did. A model that scores 96% because you copied a configuration teaches you nothing. A model that goes from 72% to 85% because you added augmentation, and you can show the training curves that prove it, teaches you a great deal.
The number at the top of that table is not the goal. Being able to attribute each rung of the climb to a specific change, with the curves to prove it, is.
What you are building
By the end you should have a repository containing a reproducible training pipeline, at least three models trained under controlled conditions, an evaluation that goes beyond a single accuracy number, and a short written analysis of what worked and what did not.
A structure that keeps this manageable:
cifar10/ data.py # dataset, transforms, loaders models.py # SimpleCNN, ResNet variants, pretrained wrapper train.py # training loop, checkpointing, logging evaluate.py # metrics, confusion matrix, error analysis experiments.py # runs the controlled comparisons configs/ # one config file per experiment results/ # curves, tables, saved predictions README.md # findingsKeeping the model definitions out of the training script matters more than it sounds. When you come to run six experiments that differ in one setting each, you want to change one line in a config, not fork the training loop six times and lose track of which version produced which number.
Look at the data before you model it
Load the dataset, plot a 10×10 grid of images with their labels, and actually look at it. This takes ten minutes and prevents a category of mistake that is otherwise invisible.
Check three things specifically. Is the class distribution balanced? CIFAR-10 has exactly 5000 training images per class, so if your loader reports otherwise, your loader is wrong. Do the labels match the images when you plot them? A transposed or off-by-one label mapping produces a model that trains to about 10% and leaves you debugging the architecture for two days. And what are the per-channel means and standard deviations of the training split?
Compute those statistics rather than copying them, because you will need them for normalisation and because computing them proves your data pipeline works:
1import torch2from torchvision import datasets3from torchvision.transforms import v245to_tensor = v2.Compose([v2.ToImage(), v2.ToDtype(torch.float32, scale=True)])6train = datasets.CIFAR10(root="./data", train=True, download=True,7 transform=to_tensor)8loader = torch.utils.data.DataLoader(train, batch_size=1000, num_workers=4)910n, mean, sq = 0, torch.zeros(3), torch.zeros(3)11for x, _ in loader:12 b = x.size(0)13 x = x.view(b, 3, -1)14 n += b15 mean += x.mean(dim=2).sum(dim=0)16 sq += (x ** 2).mean(dim=2).sum(dim=0)1718mean /= n19std = (sq / n - mean ** 2).sqrt()20print(mean, std) # roughly [0.4914, 0.4822, 0.4465], [0.2470, 0.2435, 0.2616]Note the split: statistics come from the training set only. Computing them over train plus test leaks information about the test distribution into your training pipeline, and your final number becomes slightly dishonest in a way you cannot quantify afterwards.
You also need a validation split. Carve 5,000 images out of the 50,000 training images and hold them back. The 10,000 test images are for one final measurement at the very end. If you tune hyperparameters against the test set, your test accuracy stops being an estimate of real-world performance and becomes a number you have optimised directly — which is exactly the thing it was supposed to measure independently.
The preprocessing pipeline
Two pipelines, one for training and one for evaluation, differing only in augmentation.
1import torch2from torchvision.transforms import v234MEAN = (0.4914, 0.4822, 0.4465)5STD = (0.2470, 0.2435, 0.2616)67train_tf = v2.Compose([8 v2.RandomCrop(32, padding=4), # pad then crop -> random translation9 v2.RandomHorizontalFlip(p=0.5),10 v2.ToImage(),11 v2.ToDtype(torch.float32, scale=True), # -> [0, 1]12 v2.Normalize(MEAN, STD),13 v2.RandomErasing(p=0.25), # tensor-space, so it comes last14])1516eval_tf = v2.Compose([17 v2.ToImage(),18 v2.ToDtype(torch.float32, scale=True),19 v2.Normalize(MEAN, STD),20])Three decisions in there are worth defending.
Random crop with padding, not resizing. The images are already 32×32. Padding by 4 and cropping back to 32 simulates the object sitting slightly off-centre, which is the variation you want. Scaling up and cropping would just blur.
Horizontal flip but not vertical. A mirrored car is a plausible car. An upside-down car is not something the model will ever encounter, so training on it wastes capacity and can mislead.
No colour jitter initially. Colour carries real signal in this dataset — frogs are green, ships sit on blue water. Aggressive hue shifts can push an image towards a class it does not belong to. Add it later as a controlled experiment and measure whether it helps.
Normalisation is identical in both pipelines, and that is non-negotiable. Augmentation is supposed to differ between training and evaluation; normalisation is not.
Three models, built in order
A baseline CNN you write yourself
Start with something small enough to understand completely: three convolutional blocks, each doubling the channels and halving the spatial size, then global average pooling and a linear layer.
1import torch.nn as nn23def block(c_in, c_out):4 return nn.Sequential(5 nn.Conv2d(c_in, c_out, 3, padding=1, bias=False),6 nn.BatchNorm2d(c_out),7 nn.ReLU(inplace=True),8 nn.Conv2d(c_out, c_out, 3, padding=1, bias=False),9 nn.BatchNorm2d(c_out),10 nn.ReLU(inplace=True),11 nn.MaxPool2d(2),12 )1314class SimpleCNN(nn.Module):15 def __init__(self, num_classes=10):16 super().__init__()17 self.features = nn.Sequential(18 block(3, 64), # 32 -> 1619 block(64, 128), # 16 -> 820 block(128, 256), # 8 -> 421 )22 self.head = nn.Sequential(23 nn.AdaptiveAvgPool2d(1),24 nn.Flatten(),25 nn.Dropout(0.3),26 nn.Linear(256, num_classes),27 )2829 def forward(self, x):30 return self.head(self.features(x))Two details carry weight. bias=False on convolutions followed by batch normalisation, because batch norm has its own learnable shift and the convolution bias is redundant — it would be subtracted out immediately. And global average pooling instead of flattening: flattening 4×4×256 gives 4096 values, which connected to 10 classes costs 41,000 parameters and invites overfitting, while pooling costs none.
A ResNet sized for 32×32
The standard ImageNet ResNet begins with a 7×7 convolution at stride 2 followed by a max pool, which takes a 224×224 input down to 56×56 immediately. Apply that to a 32×32 image and you are at 8×8 before the first residual block, having destroyed most of the spatial information you had.
The CIFAR variant replaces that stem with a single 3×3 convolution at stride 1 and no max pool. Everything else stays. This one change is worth several percentage points, and understanding why is worth more than the points: architectures carry assumptions about input scale, and transplanting one to a different resolution without adjusting the stem is a standard way to lose accuracy for no reason.
1from torchvision.models import resnet1823def resnet18_cifar(num_classes=10):4 m = resnet18(weights=None, num_classes=num_classes)5 m.conv1 = nn.Conv2d(3, 64, kernel_size=3, stride=1, padding=1, bias=False)6 m.maxpool = nn.Identity() # remove the aggressive early downsample7 return mA pretrained backbone, fine-tuned
Here you face a real tension. ImageNet weights were learned at 224×224. Your images are 32×32. Two options, and you should try both because the answer is not obvious in advance.
| Option | Method | Trade-off |
|---|---|---|
| Upscale the input | Resize 32→224, keep the model unchanged | Best accuracy; roughly 49× the compute per image, and the upscaling invents no real detail |
| Adapt the stem | Keep 32×32, swap the stem as above, keep pretrained weights for the rest | Fast; the pretrained stem is discarded and the remaining layers see feature scales they were not trained on |
Upscaling usually wins on accuracy despite being obviously wasteful, which is itself an instructive result about how much the pretrained features are worth.
Whichever you pick, follow the two-phase schedule. Freeze the backbone and train only the new final layer for two or three epochs, then unfreeze and continue at a learning rate ten to a hundred times lower. Skipping the warm-up means the randomly initialised head produces large meaningless gradients that flow back and damage the pretrained weights in the first few batches — the exact thing you were trying to preserve.
Training
A configuration that works reliably as a starting point:
| Setting | Value | Reasoning |
|---|---|---|
| Optimiser | SGD, momentum 0.9 | Generalises slightly better than Adam on this dataset; AdamW is a fine alternative |
| Learning rate | 0.1 from scratch, 0.001 fine-tuning | Refining good weights needs far smaller steps than searching from random ones |
| Schedule | Cosine decay, 5-epoch warm-up | Warm-up prevents early divergence at high learning rates |
| Weight decay | 5e-4 | Standard for CIFAR; do not apply it to batch norm parameters |
| Batch size | 128 | Fits comfortably; scale the learning rate proportionally if you change it |
| Epochs | 50–100 from scratch, 10–20 fine-tuning | Watch validation, not the epoch counter |
| Label smoothing | 0.1 | Discourages overconfidence and reliably adds a fraction of a point |
Log training loss, validation loss and validation accuracy every epoch, and save a checkpoint whenever validation accuracy improves. The final epoch is frequently not the best one, and without checkpointing you will discover that fact and have nothing to show for it.
Set every random seed and record it. A run you cannot repeat is not a result, and the difference between two configurations is often smaller than the difference between two seeds of the same configuration — which is exactly why you need to know which is which.
Evaluate properly
A single accuracy figure is the least interesting thing you can produce. Four things tell you far more.
Per-class accuracy. Overall accuracy of 88% might hide 96% on ships and 71% on cats. That imbalance is a finding, and it points somewhere specific.
The confusion matrix. This is where the real story lives. On CIFAR-10 the errors are highly structured and consistent across models: cat and dog are confused constantly, deer and horse frequently, aeroplane and ship occasionally. The last of those is the interesting one — both are grey shapes against a uniform blue background, and at 32×32 that background dominates. Your model has partly learned to classify backgrounds.
The actual misclassified images. Plot fifty of them with their true and predicted labels. You will find genuinely ambiguous images, a few plainly wrong labels, and clusters that reveal a systematic weakness. This is the highest-value ten minutes in the entire project.
Calibration. Bucket predictions by confidence and check whether accuracy within each bucket matches it. If the predictions with 90% confidence are right only 70% of the time, the model is overconfident, which matters enormously for anything that acts on the output rather than just reporting it.
Run controlled experiments
This is the part that separates a project from a tutorial. Change one thing at a time, keep everything else fixed including the seed, and record the result in a table.
| Experiment | Vary | Question it answers |
|---|---|---|
| Augmentation ablation | None / flip only / flip + crop / full | How much does each augmentation contribute? |
| Batch norm ablation | With and without | Does it improve accuracy, or only training speed? |
| Depth | ResNet-18 / 34 / 50 | Where does depth stop paying off on 32×32 inputs? |
| Optimiser | SGD+momentum vs AdamW | Does the folklore hold on your setup? |
| Data volume | 10% / 25% / 50% / 100% of training data | Would more data help, or has it plateaued? |
| Transfer | Scratch / frozen features / fine-tuned | What are pretrained features actually worth here? |
Change one thing at a time, or you will end up with a better model and no idea which of your five changes produced it.
The data-volume experiment is the most useful and the most often skipped. Plot accuracy against training-set size on a log scale. If the curve is still climbing steeply at 100%, more data is your cheapest improvement. If it has flattened, more data will not help and you should be changing the model or the training recipe instead. Teams routinely spend months collecting data that a two-hour experiment would have shown to be useless.
When you report results, report the gap between training and validation accuracy alongside each number. Two configurations reaching 89% mean different things if one has a 2-point gap and the other has a 9-point gap — the second is overfitting and will degrade faster on genuinely new data.
What to take away from it
The deliverable that matters is not the highest number you reached. It is the table of experiments with one variable changed per row, and a paragraph of honest interpretation next to it. Anyone can download a configuration that scores 96%. Being able to say "augmentation was worth 11 points, batch normalisation was worth 4, and pretraining was worth 6, and here is the evidence" is a different and much more transferable skill.
Be honest about the negative results too. If AdamW performed worse than SGD on your setup, write that down with the numbers rather than quietly dropping it. If a change you were confident about made no difference, that is a real finding about how much that technique matters at this scale, and it will stop you from cargo-culting it into the next project.
Finally, notice which of your accuracy gains came from the model and which came from everything around it. On this dataset, augmentation and training schedule typically contribute more than architecture does. That ratio holds in most real applications too, and it is the single most useful thing you can carry forward: when a model underperforms, the data pipeline and the training recipe are usually where the points are hiding, not the layer diagram.