- MantraMindAI
- Courses
- Multimodal AI
- Multimodal Vision-Language Models
Multimodal Vision-Language Models
Course Overview
Intermediate
Free Course
For engineers who know the basics of deep learning and want to build products that combine images and text. You will be able to use CLIP-style embeddings for search and tagging, run generative vision-language models for question answering and captioning, and ship a working multimodal image assistant.
Instructor: Jaidev
Sections: 3
Course Content
Section 1: Vision-Language Integration
How CLIP aligns images and text in one vector space, and how the LLaVA and Flamingo designs let a language model see. 2 lessons, about 35 minutes.
Section 2: Applications
Build visual question answering, captioning and cross-modal search that hold up in production, with the metrics to prove it. 2 lessons, about 30 minutes.
Section 3: Mini Project
Build a photo search and question-answering assistant from CLIP embeddings and a generative vision-language model. 1 lesson, about 15 minutes.