- MantraMindAI
- Courses
- Multimodal AI
- Computer Vision Fundamentals
Computer Vision Fundamentals
For developers who know basic Python and want to understand how image models work, from pixels up to convolutional networks. You will be able to preprocess and augment images correctly, build and fine-tune CNNs in PyTorch, and train and evaluate your own image classifier.
Course Content
How a computer stores an image as numbers, and how to resize, normalise and augment images without silently breaking a model. 3 lessons, about 45 minutes.
How convolution and pooling work, how the classic architectures from LeNet to ResNet fixed real training failures, and how to reuse pretrained networks for classification, detection and segmentation. 4 lessons, about 50 minutes.
Train and compare CNNs on CIFAR-10 step by step, and attribute each gain in accuracy to a specific change. 1 lesson, about 15 minutes.