M24.6 CONNECT THE MECHANISM
Learn a policy directly from its outcomes
Skip the value table and tune the odds of each move directly. Meet REINFORCE, actor-critic, and PPO, the algorithm that helped turn language models into chat assistants.
LESSON OVERVIEW13 min lesson
Lesson overview
Skip the value table and tune the odds of each move directly. Meet REINFORCE, actor-critic, and PPO, the algorithm that helped turn language models into chat assistants.
What you’ll explore
- Policy-gradient methods adjust action probabilities using return-related signals; actor-critic methods learn value baselines, and PPO-style objectives control update incentives rather than guaranteeing optimality.
GO TO THE SOURCE
Original explanations, connected to the research.
Simple statistical gradient-following algorithms for connectionist reinforcement learning (Williams, 1992)Policy Gradient Methods for Reinforcement Learning with Function Approximation (Sutton et al., 1999)Proximal Policy Optimization Algorithms (Schulman et al., 2017)Training language models to follow instructions with human feedback (Ouyang et al., 2022)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.