M15.6 CONNECT THE MECHANISM
Separate mixing across tokens from transforming one token
Every transformer layer repeats one small recipe: normalize, attend, add; normalize, MLP, add. Trace it with real numbers and count the 7 million weights in each GPT-2 layer.
LESSON OVERVIEW12 min lesson
Lesson overview
Every transformer layer repeats one small recipe: normalize, attend, add; normalize, MLP, add. Trace it with real numbers and count the 7 million weights in each GPT-2 layer.
What you’ll explore
- Transformer blocks combine attention, position-wise feed-forward computation, residual additions, and normalization; gated activations and placement conventions change the exact computation.
GO TO THE SOURCE
Original explanations, connected to the research.
Attention Is All You Need, sections 3.1 and 3.3 (Vaswani et al., 2017)Language Models are Unsupervised Multitask Learners (GPT-2; Radford et al., 2019)On Layer Normalization in the Transformer Architecture (Xiong et al., 2020)Root Mean Square Layer Normalization (Zhang & Sennrich, 2019)GLU Variants Improve Transformer (Shazeer, 2020)A Mathematical Framework for Transformer Circuits (Elhage et al., 2021)LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023)The Llama 3 Herd of ModelsSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.