M16.4 CONNECT THE MECHANISM
Assemble token batches without leaking targets
Pad every document to 2,048 tokens and most of each batch is blank. Pack them end to end instead, then decide which tokens may look at one another.
LESSON OVERVIEW14 min lesson
Lesson overview
Pad every document to 2,048 tokens and most of each batch is blank. Pack them end to end instead, then decide which tokens may look at one another.
What you’ll explore
- Distinguish padding, attention masks, and loss masks using shifted next-token targets and document boundaries.
GO TO THE SOURCE
Original explanations, connected to the research.
Brown et al. — Language Models are Few-Shot Learners (GPT-3, 2020), section 2.3 and appendix BThe Llama 3 Herd of Models (2024), section 3.4Groeneveld et al. — OLMo, an open language-model frameworkDive into Deep Learning — computational graphs and backpropagationSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.