Reinforcement Learning from Human Feedback (RLHF)

Course Overview
Advanced
Free Course

For ML engineers who already fine-tune language models and now need to align them with human preferences. You will be able to build preference datasets and reward models, and train and evaluate policies with PPO, DPO and its variants, GRPO-style RL with verifiable rewards, and constitutional methods using current open-source tooling.

Instructor: Jaidev
Sections: 4

Course Content

Section 1: Reward Modeling Fundamentals

Why preference learning exists, how human comparisons become a reward model, and how PPO optimises a policy against that reward without drifting off a cliff.

Section 2: Training Workflow

How to collect preference data you can trust, train and audit a reward model, run PPO safely with today's tools, and turn safety goals into data and evaluations.

Section 3: Alternative Techniques

The methods that replace or complement PPO today: DPO and its variants, constitutional AI and AI feedback, GRPO-style RL with verifiable rewards, and the open-source frameworks that run them.

Section 4: Mini Project

An end-to-end build you run yourself: real preference pairs from OASST1, an audited reward model, a DPO-aligned policy, and an evaluation that can catch your own mistakes.