M16.7 CONNECT THE MECHANISM
Allocate compute and preserve useful checkpoints
Six times parameters times tokens. One multiplication prices Kestrel's whole run in GPU-days, and Chinchilla says it should be twice as big. Here's why the team ignores that.
LESSON OVERVIEW15 min lesson
Lesson overview
Six times parameters times tokens. One multiplication prices Kestrel's whole run in GPU-days, and Chinchilla says it should be twice as big. Here's why the team ignores that.
What you’ll explore
- Pretraining budgets trade model size, token exposure, and runtime; scaling trends guide experiments while checkpoint evaluation detects capability changes and failures.
GO TO THE SOURCE
Original explanations, connected to the research.
Hoffmann et al. — Training Compute-Optimal Large Language Models (Chinchilla, 2022)Kaplan et al. — Scaling Laws for Neural Language Models (2020)Touvron et al. — LLaMA: Open and Efficient Foundation Language Models (2023)Chowdhery et al. — PaLM: Scaling Language Modeling with Pathways (2022), which defines model FLOPs utilizationBiderman et al. — Pythia, a suite for analyzing language-model training (2023)The Llama 3 Herd of Models (2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.