M15.3 CONNECT THE MECHANISM
From attention scores to a causal mask
One line of math, softmax(QKᵀ/√d_k)V, runs inside every transformer. Work it by hand on three tokens, and see why one missing mask lets a model cheat.
LESSON OVERVIEW17 min lesson
Lesson overview
One line of math, softmax(QKᵀ/√d_k)V, runs inside every transformer. Work it by hand on three tokens, and see why one missing mask lets a model cheat.
What you’ll explore
- Compute normalized attention weights and explain how a causal mask prevents future-token leakage.
GO TO THE SOURCE
Original explanations, connected to the research.
Attention Is All You Need, section 3.2.1 and footnote 4 (Vaswani et al., 2017)Dive into Deep Learning — attention scoring functionsThe Annotated Transformer (Harvard NLP)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.