M21.8 CONNECT THE MECHANISM
Transfer reasoning behavior and compare its operating cost
A 32-billion-parameter model copied a bigger model's worked solutions and beat the same model trained with RL by 25 points. When is distillation worth it, and what does a solved problem really cost?
LESSON OVERVIEW12 min lesson
Lesson overview
A 32-billion-parameter model copied a bigger model's worked solutions and beat the same model trained with RL by 25 points. When is distillation worth it, and what does a solved problem really cost?
What you’ll explore
- Reasoning distillation trains students on selected teacher behavior; compare accuracy, explanation quality, output length, latency, and training provenance at matched budgets.
GO TO THE SOURCE
Original explanations, connected to the research.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025), sections 2.4 and 4.1Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes (Hsieh et al., 2023)s1: Simple test-time scaling (Muennighoff et al., 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.