M20.1 CONNECT THE MECHANISM
Track what each adaptation stage changes
Ask a raw base model "Do you fix e-bikes?" and it may write a whole forum thread. Meet the training stages that turn it into an assistant, and what each can break.
LESSON OVERVIEW15 min lesson
Lesson overview
Ask a raw base model "Do you fix e-bikes?" and it may write a whole forum thread. Meet the training stages that turn it into an assistant, and what each can break.
What you’ll explore
- Classify supervised demonstrations, preference training, reward-based updates, and runtime-only changes by the data and parameters involved.
GO TO THE SOURCE
Original explanations, connected to the research.
Training language models to follow instructions with human feedback (InstructGPT; Ouyang et al., 2022)Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al., 2023)Tülu 3: Pushing Frontiers in Open Language Model Post-Training (Lambert et al., 2024)DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.