M19.3 CONNECT THE MECHANISM
Rearrange attention around a compressed running summary
What if a transformer kept one running total instead of a list of every past token? Move one pair of brackets and attention turns into an RNN, with a catch.
LESSON OVERVIEW14 min lesson
Lesson overview
What if a transformer kept one running total instead of a list of every past token? Move one pair of brackets and attention turns into an RNN, with a catch.
What you’ll explore
- Linear-attention methods replace or approximate the attention kernel so sums can be reassociated; they reduce sequence-length costs under assumptions while changing the representation and its limits.
GO TO THE SOURCE
Original explanations, connected to the research.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention (Katharopoulos et al., 2020)Rethinking Attention with Performers (Choromanski et al., 2020)Repeat After Me: Transformers are Better than State Space Models at Copying (Jelassi et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.