M19.2 CONNECT THE MECHANISM
Reduce the state stored for attention keys and values
Sixteen users chatting with a 70B model can need 160 GiB of attention cache, more than the model itself. Learn the one-line formula, and the tricks that cut it 8× and beyond.
LESSON OVERVIEW14 min lesson
Lesson overview
Sixteen users chatting with a 70B model can need 160 GiB of attention cache, more than the model itself. Learn the one-line formula, and the tricks that cut it 8× and beyond.
What you’ll explore
- Multi-query and grouped-query attention share key/value heads, while latent attention compresses their representation; cache savings depend on exact dimensions and implementation.
GO TO THE SOURCE
Original explanations, connected to the research.
Fast Transformer Decoding: One Write-Head is All You Need (Shazeer, 2019)GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (Ainslie et al., 2023)Llama 2: Open Foundation and Fine-Tuned Chat Models (Touvron et al., 2023)DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (DeepSeek-AI, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.