M21.2 CONNECT THE MECHANISM
Teach a useful starting style for problem solving
A model trained purely by reward learned to reason but mixed languages mid-thought. A few thousand good worked examples fixed that. Learn what makes a demonstration worth copying.
LESSON OVERVIEW14 min lesson
Lesson overview
A model trained purely by reward learned to reason but mixed languages mid-thought. A few thousand good worked examples fixed that. Learn what makes a demonstration worth copying.
What you’ll explore
- Reasoning demonstrations and cold-start tuning supply examples of problem decomposition and presentation; curate valid, diverse traces and distinguish supervised imitation from later reward optimization.
GO TO THE SOURCE
Original explanations, connected to the research.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025)STaR: Bootstrapping Reasoning With Reasoning (Zelikman et al., 2022)Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes (Hsieh et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.