M17.7 CONNECT THE MECHANISM
Compute attention with less memory traffic
Standard attention writes a gigabyte-sized score matrix to memory just to read it back. FlashAttention never writes it, and still gets exactly the same answer. Here's the trick.
LESSON OVERVIEW14 min lesson
Lesson overview
Standard attention writes a gigabyte-sized score matrix to memory just to read it back. FlashAttention never writes it, and still gets exactly the same answer. Here's the trick.
What you’ll explore
- Tiled attention kernels can preserve dense attention mathematics while avoiding a fully stored score matrix; kernel fusion and overlap improve execution without automatically changing asymptotic arithmetic.
GO TO THE SOURCE
Original explanations, connected to the research.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)Online normalizer calculation for softmax (Milakov & Gimelshein, 2018)FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (Dao, 2023)FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (Shah et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.