Back to the lesson libraryMECHANISM · 14 MIN
M21.4 CONNECT THE MECHANISM

Compare several sampled solutions to form a baseline

Ask the same maths question eight times, and the answers grade each other. That one idea, GRPO, removed a whole network from RL training and powered DeepSeek-R1.

LESSON OVERVIEW14 min lesson

Lesson overview

Ask the same maths question eight times, and the answers grade each other. That one idea, GRPO, removed a whole network from RL training and powered DeepSeek-R1.

What you’ll explore

  • GRPO scores each answer against the other answers to the same prompt, then updates the policy with clipping and a KL penalty; the group's spread and sampling choices shape the signal.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.