M15.4 CONNECT THE MECHANISM
Run several learned comparisons in parallel
GPT-2 small splits every attention layer into 12 heads of 64 numbers each. Trace the split, the concat, and the output projection, and count the 2.4 million weights involved.
LESSON OVERVIEW10 min lesson
Lesson overview
GPT-2 small splits every attention layer into 12 heads of 64 numbers each. Trace the split, the concat, and the output projection, and count the 2.4 million weights involved.
What you’ll explore
- Compute per-head attention outputs, concatenate them, and apply the output projection; track head width (d_model = h·d_head) and projection parameters (4·d_model²) without treating heads as guaranteed named specialists.
GO TO THE SOURCE
Original explanations, connected to the research.
Attention Is All You Need, section 3.2.2 (Vaswani et al., 2017)Language Models are Unsupervised Multitask Learners (GPT-2; Radford et al., 2019)What Does BERT Look At? An Analysis of BERT's Attention (Clark et al., 2019)Are Sixteen Heads Really Better than One? (Michel, Levy & Neubig, 2019)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.