M21.3 CONNECT THE MECHANISM
Turn checkable outcomes into a learning signal
For maths and code you don't need a judge; you can just check the answer. So why did a model being trained on coding tasks learn to call exit(0)?
LESSON OVERVIEW13 min lesson
Lesson overview
For maths and code you don't need a judge; you can just check the answer. So why did a model being trained on coding tasks learn to call exit(0)?
What you’ll explore
- Verifiable rewards score outputs with explicit checks such as mathematical equivalence or program tests; verifier coverage and isolation determine how closely reward tracks the intended task.
GO TO THE SOURCE
Original explanations, connected to the research.
Tülu 3: Pushing Frontiers in Open Language Model Post-Training (Lambert et al., 2024)DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025)Training Verifiers to Solve Math Word Problems (Cobbe et al., 2021)Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (Baker et al., 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.