M17.2 CONNECT THE MECHANISM
Account for everything that occupies accelerator memory
A 7B model's weights fill just 14 GB, yet training it won't fit on an 80 GB GPU. Find the other 100+ gigabytes, line by line.
LESSON OVERVIEW14 min lesson
Lesson overview
A 7B model's weights fill just 14 GB, yet training it won't fit on an 80 GB GPU. Find the other 100+ gigabytes, line by line.
What you’ll explore
- Training memory includes parameters, gradients, optimizer state, activations, buffers, and allocator overhead; inference has a different budget including its key/value cache.
GO TO THE SOURCE
Original explanations, connected to the research.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models (Rajbhandari et al., 2019)Reducing Activation Recomputation in Large Transformer Models (Korthikanti et al., 2022), section 4.1Mixed Precision Training (Micikevicius et al., 2017)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.