Back to the lesson libraryMECHANISM · 14 MIN
M19.2 CONNECT THE MECHANISM

Reduce the state stored for attention keys and values

Sixteen users chatting with a 70B model can need 160 GiB of attention cache, more than the model itself. Learn the one-line formula, and the tricks that cut it 8× and beyond.

LESSON OVERVIEW14 min lesson

Lesson overview

Sixteen users chatting with a 70B model can need 160 GiB of attention cache, more than the model itself. Learn the one-line formula, and the tricks that cut it 8× and beyond.

What you’ll explore

  • Multi-query and grouped-query attention share key/value heads, while latent attention compresses their representation; cache savings depend on exact dimensions and implementation.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.