M20.6 CONNECT THE MECHANISM
Train on preference pairs without a separate online reward loop
RLHF juggles four models and a sampling loop. DPO gets the same kind of result from one line of loss and two models, because the assistant turns out to be its own reward model.
LESSON OVERVIEW13 min lesson
Lesson overview
RLHF juggles four models and a sampling loop. DPO gets the same kind of result from one line of loss and two models, because the assistant turns out to be its own reward model.
What you’ll explore
- DPO adjusts chosen-versus-rejected response likelihoods relative to a reference model using a preference loss; its assumptions and data coverage still constrain resulting behavior.
GO TO THE SOURCE
Original explanations, connected to the research.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al., 2023)Zephyr: Direct Distillation of LM Alignment (Tunstall et al., 2023)The Llama 3 Herd of Models (Llama Team, Meta, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.