M21.4 CONNECT THE MECHANISM
Compare several sampled solutions to form a baseline
Ask the same maths question eight times, and the answers grade each other. That one idea, GRPO, removed a whole network from RL training and powered DeepSeek-R1.
LESSON OVERVIEW14 min lesson
Lesson overview
Ask the same maths question eight times, and the answers grade each other. That one idea, GRPO, removed a whole network from RL training and powered DeepSeek-R1.
What you’ll explore
- GRPO scores each answer against the other answers to the same prompt, then updates the policy with clipping and a KL penalty; the group's spread and sampling choices shape the signal.
GO TO THE SOURCE
Original explanations, connected to the research.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., 2024), section 4DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025)DAPO: An Open-Source LLM Reinforcement Learning System at Scale (Yu et al., 2025)Understanding R1-Zero-Like Training: A Critical Perspective (Liu et al., 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.