M09.7 CONNECT THE MECHANISM
More capacity changes more than one thing
Make a language model ten times bigger and its loss falls by a predictable amount. Learn to read scaling laws, see why test error can rise and then fall again, and meet memorization.
LESSON OVERVIEW14 min lesson
Lesson overview
Make a language model ten times bigger and its loss falls by a predictable amount. Learn to read scaling laws, see why test error can rise and then fall again, and meet memorization.
What you’ll explore
- Read scaling curves, double descent, and memorization as results from specific setups, not a law that bigger is always better.
GO TO THE SOURCE
Original explanations, connected to the research.
Scaling Laws for Neural Language Models (Kaplan et al., 2020)Training Compute-Optimal Large Language Models (Hoffmann et al., 2022)Reconciling modern machine-learning practice and the bias-variance trade-off (Belkin et al., 2019)Extracting Training Data from Large Language Models (Carlini et al., 2021)Quantifying Memorization Across Neural Language Models (Carlini et al., 2022)Deep Double Descent: Where Bigger Models and More Data Hurt (Nakkiran et al., 2019)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.