- MantraMindAI
- Courses
- Multimodal AI
- Image and Video Generation
Image and Video Generation
For engineers who want to understand how modern image and video generators work, from the maths of diffusion to the tools that control and fine-tune them. You will be able to run, steer and adapt Stable Diffusion pipelines in code, judge video generation platforms on your own shots, and build a small tool that produces a consistent image series.
Course Content
How latent diffusion turns noise into an image, worked through in real numbers, and how the denoising network is built and trained. 2 lessons, about 40 minutes.
Take control of where and what diffusion paints with ControlNet, inpainting, outpainting and image-to-image, then teach a model new subjects and styles with LoRA, DreamBooth and textual inversion. 2 lessons, about 30 minutes.
Fill in missing frames with optical flow and learned interpolation, and choose and use commercial video generators by the trade-offs that actually matter for your work. 2 lessons, about 35 minutes.
Build a tool that generates a coherent image series, with a pinned style, one named axis of variation and a number that measures consistency. 1 lesson, about 20 minutes.