Evaluating and Testing GenAI Models

Course Overview
Intermediate
Free Course

For engineers who ship generative AI features and need to know whether a model or prompt change actually made things better. You will be able to choose and compute the right metrics, detect and measure hallucinations, run human and LLM-judge evaluations, and report every result with the uncertainty it deserves.

Instructor: Jaidev
Sections: 4

Course Content

Section 1: Evaluation Metrics for Generative Models

The metrics for scoring generated text — BLEU, ROUGE, METEOR, BERTScore, perplexity and rubric scores — what each one measures, and the statistics that tell you whether a difference is real. 4 lessons, about 60 minutes.

Section 2: Detecting and Measuring Hallucinations

How hallucinations arise, how to catch them with benchmarks, fact-checking pipelines and entailment checks, and how to report them by type and severity. 4 lessons, about 55 minutes.

Section 3: Human and Hybrid Evaluation Systems

Design human evaluations that produce trustworthy numbers, combine them with validated LLM judges, and build dashboards that show their own uncertainty. 4 lessons, about 60 minutes.

Section 4: Mini Project

Build a complete pipeline that compares two models and write a recommendation the evidence can support. 1 lesson, about 20 minutes.