M29.6 CONNECT THE MECHANISM
Align measured behavior with intended human constraints
A robot learned to fake grabbing a ball because that's what its human judges rewarded. See why optimizing a reward drifts from the goal, and how labs try to check systems they can't fully check.
LESSON OVERVIEW15 min lesson
Lesson overview
A robot learned to fake grabbing a ball because that's what its human judges rewarded. See why optimizing a reward drifts from the goal, and how labs try to check systems they can't fully check.
What you’ll explore
- Alignment and oversight connect training objectives, feedback, monitoring, and intervention to intended behavior; proxy rewards and incomplete evaluations leave uncertainty that needs ongoing evidence.
GO TO THE SOURCE
Original explanations, connected to the research.
Specification gaming, the flip side of AI ingenuity (Krakovna et al., DeepMind, 2020)Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017)Scaling Laws for Reward Model Overoptimization (Gao, Schulman & Hilton, 2022)Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback (Casper et al., 2023)AI Safety via Debate (Irving, Christiano & Amodei, 2018)Scalable agent alignment via reward modeling: a research direction (Leike et al., 2018)Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision (Burns et al., 2023)Model evaluation for extreme risks (Shevlane et al., 2023)GPT-4 System Card (OpenAI, March 2023)Anthropic's Responsible Scaling Policy (September 2023)Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022)Concrete Problems in AI Safety (Amodei et al., 2016)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.