M20.4 CONNECT THE MECHANISM
Learn a scoring signal from response comparisons
Nobody can say an answer is worth "7.3". But people can say which of two answers is better, and one small formula turns thousands of those choices into a score.
LESSON OVERVIEW14 min lesson
Lesson overview
Nobody can say an answer is worth "7.3". But people can say which of two answers is better, and one small formula turns thousands of those choices into a score.
What you’ll explore
- Reward models learn to predict which answer people prefer; disagreement, biased samples, and exploitation mean a high reward score isn't proof of a good answer.
GO TO THE SOURCE
Original explanations, connected to the research.
Training language models to follow instructions with human feedback (InstructGPT; Ouyang et al., 2022)Learning to summarize from human feedback (Stiennon et al., 2020)Deep reinforcement learning from human preferences (Christiano et al., 2017)Rank analysis of incomplete block designs: I. The method of paired comparisons (Bradley & Terry, Biometrika, 1952)A Long Way to Go: Investigating Length Correlations in RLHF (Singhal et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.