M16.1 CONNECT THE MECHANISM
Choose the model that pretraining will fit
Sixteen layers, width 2,048, a 50,304-token vocabulary. See how a few numbers on a card fix a model's billion parameters before it reads a single word.
LESSON OVERVIEW15 min lesson
Lesson overview
Sixteen layers, width 2,048, a 50,304-token vocabulary. See how a few numbers on a card fix a model's billion parameters before it reads a single word.
What you’ll explore
- Explain how architecture, tokenizer, initialization, objective, and budget jointly define a base-model training experiment.
GO TO THE SOURCE
Original explanations, connected to the research.
Vaswani et al. — Attention Is All You Need (2017)Kaplan et al. — Scaling Laws for Neural Language Models (2020)Biderman et al. — Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling (2023)Groeneveld et al. — OLMo: Accelerating the Science of Language Models (2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.