Synthetic Data Generation

Course Overview
Intermediate
Free Course

For ML and AI engineers who are short of labelled, rare or privacy-restricted data. You will generate synthetic text, dialogue and image data with LLMs and diffusion models, filter it, and prove with fidelity, diversity, privacy and real-data tests whether it actually helps your model.

Instructor: Jaidev
Sections: 3

Course Content

Section 1: Synthetic Data for AI Systems

When synthetic data is worth making, how LLMs and diffusion models generate it, and how to measure its fidelity, diversity, bias and privacy before you trust it. 3 lessons, about 45 minutes.

Section 2: Techniques for Generating Synthetic Data

The working techniques: prompt-based generation with a filtering funnel, augmentation that does not corrupt labels, and multi-turn dialogue synthesis that stays consistent. 3 lessons, about 40 minutes.

Section 3: Mini Project

Build a synthetic support-dialogue pipeline end to end and prove, on real tickets, that it makes a classifier better. 1 lesson, about 20 minutes.