Multimodal Vision-Language Models

Course Overview
Intermediate
Free Course

For engineers who know the basics of deep learning and want to build products that combine images and text. You will be able to use CLIP-style embeddings for search and tagging, run generative vision-language models for question answering and captioning, and ship a working multimodal image assistant.

Instructor: Jaidev
Sections: 3

Course Content

Section 1: Vision-Language Integration

How CLIP aligns images and text in one vector space, and how the LLaVA and Flamingo designs let a language model see. 2 lessons, about 35 minutes.

Section 2: Applications

Build visual question answering, captioning and cross-modal search that hold up in production, with the metrics to prove it. 2 lessons, about 30 minutes.

Section 3: Mini Project

Build a photo search and question-answering assistant from CLIP embeddings and a generative vision-language model. 1 lesson, about 15 minutes.