M19.1 CONNECT THE MECHANISM
Choose which token pairs may communicate
A 100,000-token codebase means five billion token pairs per head, per layer. See how sliding windows, dilation, and global tokens cut that bill, and what each cut gives up.
LESSON OVERVIEW16 min lesson
Lesson overview
A 100,000-token codebase means five billion token pairs per head, per layer. See how sliding windows, dilation, and global tokens cut that bill, and what each cut gives up.
What you’ll explore
- Local, sliding-window, sparse, and global attention patterns change direct token connectivity; trace multi-layer information paths and measure quality alongside reduced pair counts.
GO TO THE SOURCE
Original explanations, connected to the research.
Longformer: The Long-Document Transformer (Beltagy, Peters & Cohan, 2020)Big Bird: Transformers for Longer Sequences (Zaheer et al., 2020)Generating Long Sequences with Sparse Transformers (Child et al., 2019)Mistral 7B (Jiang et al., 2023)Gemma 2: Improving Open Language Models at a Practical Size (Gemma Team, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.