M17.1 CONNECT THE MECHANISM
Why arithmetic counts do not tell the whole speed story
On a GPU, pushing 64 tokens through a weight matrix takes about as long as pushing one. Learn to count bytes as well as FLOPs, and read the roofline.
LESSON OVERVIEW14 min lesson
Lesson overview
On a GPU, pushing 64 tokens through a weight matrix takes about as long as pushing one. Learn to count bytes as well as FLOPs, and read the roofline.
What you’ll explore
- Runtime depends on arithmetic throughput, memory traffic, communication, and scheduling; arithmetic intensity helps identify which resource limits an operation.
GO TO THE SOURCE
Original explanations, connected to the research.
Roofline: An Insightful Visual Performance Model for Multicore Architectures (Williams, Waterman & Patterson, Communications of the ACM, 2009)NVIDIA A100 Tensor Core GPU datasheetPaLM: Scaling Language Modeling with Pathways (Chowdhery et al., 2022), which introduced model FLOPs utilizationLlama 2: Open Foundation and Fine-Tuned Chat Models (Touvron et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.