M35.5 CONNECT THE MECHANISM
Separate DeepSeek-R1, R1-Zero, and distilled students
Download "DeepSeek-R1-Distill-Qwen-32B" and you get a Qwen model, not a small DeepSeek. Untangle R1-Zero, R1, and their students, and find out why the student caches more per token than its giant teacher.
LESSON OVERVIEW14 min lesson
Lesson overview
Download "DeepSeek-R1-Distill-Qwen-32B" and you get a Qwen model, not a small DeepSeek. Untangle R1-Zero, R1, and their students, and find out why the student caches more per token than its giant teacher.
What you’ll explore
- DeepSeek-R1 illustrates reasoning post-training on an already pretrained base; distinguish the staged R1 recipe, direct-RL R1-Zero experiment, and architecturally different distilled students.
GO TO THE SOURCE
Original explanations, connected to the research.
DeepSeek-R1 report v1DeepSeek-R1 official repositorySuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.