M16.5 CONNECT THE MECHANISM
Choose what a pretraining prediction means
GPT guesses the next word, BERT fills in blanks, and T5 rebuilds missing phrases. Work out Kestrel's loss by hand and see why one game won.
LESSON OVERVIEW13 min lesson
Lesson overview
GPT guesses the next word, BERT fills in blanks, and T5 rebuilds missing phrases. Work out Kestrel's loss by hand and see why one game won.
What you’ll explore
- Trace causal next-token targets and distinguish them from masked-token and reconstruction objectives, including loss masking and reduction.
GO TO THE SOURCE
Original explanations, connected to the research.
Devlin et al. — BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018)Raffel et al. — Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5, 2019)Brown et al. — Language Models are Few-Shot Learners (GPT-3, 2020)Biderman et al. — Pythia, a suite for analyzing language-model trainingSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.