Back to the lesson libraryMECHANISM · 14 MIN
M19.3 CONNECT THE MECHANISM

Rearrange attention around a compressed running summary

What if a transformer kept one running total instead of a list of every past token? Move one pair of brackets and attention turns into an RNN, with a catch.

LESSON OVERVIEW14 min lesson

Lesson overview

What if a transformer kept one running total instead of a list of every past token? Move one pair of brackets and attention turns into an RNN, with a catch.

What you’ll explore

  • Linear-attention methods replace or approximate the attention kernel so sums can be reassociated; they reduce sequence-length costs under assumptions while changing the representation and its limits.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.