- MantraMindAI
- Courses
- AI Engineering & MLOps
- Evaluating and Testing GenAI Models
Evaluating and Testing GenAI Models
For engineers who ship generative AI features and need to know whether a model or prompt change actually made things better. You will be able to choose and compute the right metrics, detect and measure hallucinations, run human and LLM-judge evaluations, and report every result with the uncertainty it deserves.
Course Content
The metrics for scoring generated text — BLEU, ROUGE, METEOR, BERTScore, perplexity and rubric scores — what each one measures, and the statistics that tell you whether a difference is real. 4 lessons, about 60 minutes.
How hallucinations arise, how to catch them with benchmarks, fact-checking pipelines and entailment checks, and how to report them by type and severity. 4 lessons, about 55 minutes.
Design human evaluations that produce trustworthy numbers, combine them with validated LLM judges, and build dashboards that show their own uncertainty. 4 lessons, about 60 minutes.
Build a complete pipeline that compares two models and write a recommendation the evidence can support. 1 lesson, about 20 minutes.