M16.6 CONNECT THE MECHANISM
Inside a language model training step
Every 1.6 seconds, Kestrel reads a million tokens and nudges a billion weights. Follow one training step from batch to update, and run a tiny one yourself.
LESSON OVERVIEW14 min lesson
Lesson overview
Every 1.6 seconds, Kestrel reads a million tokens and nudges a billion weights. Follow one training step from batch to update, and run a tiny one yourself.
What you’ll explore
- Distinguish forward evaluation, gradient computation, and optimizer updates in next-token pretraining.
GO TO THE SOURCE
Original explanations, connected to the research.
Brown et al. — Language Models are Few-Shot Learners (GPT-3, 2020), appendix BBiderman et al. — Pythia, a suite for analyzing language-model training (2023)PyTorch — optimizing model parametersDive into Deep Learning — transformer architectureSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.