M20.5 CONNECT THE MECHANISM
Learn from feedback while controlling policy changes
Train an assistant to please a reward model and it soon opens every reply with "Absolutely! Great question!" One penalty term, charged token by token, keeps it honest.
LESSON OVERVIEW14 min lesson
Lesson overview
Train an assistant to please a reward model and it soon opens every reply with "Absolutely! Great question!" One penalty term, charged token by token, keeps it honest.
What you’ll explore
- An RLHF loop samples responses, scores them, estimates learning signals, and updates the policy; PPO-style constraints and reference-policy penalties address different forms of update control.
GO TO THE SOURCE
Original explanations, connected to the research.
Training language models to follow instructions with human feedback (InstructGPT; Ouyang et al., 2022)Learning to summarize from human feedback (Stiennon et al., 2020)Proximal Policy Optimization Algorithms (Schulman et al., 2017)Scaling Laws for Reward Model Overoptimization (Gao, Schulman & Hilton, 2022)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.