Back to the lesson libraryMECHANISM · 14 MIN
M17.5 CONNECT THE MECHANISM

Save memory by sharing state or recreating it

Plain data parallelism keeps 256 identical copies of the optimizer. Remove the duplicates and the 7B model's states drop from 112 GB to under half a gigabyte per GPU.

LESSON OVERVIEW14 min lesson

Lesson overview

Plain data parallelism keeps 256 identical copies of the optimizer. Remove the duplicates and the 7B model's states drop from 112 GB to under half a gigabyte per GPU.

What you’ll explore

  • State sharding, activation checkpointing, and offloading reduce different memory allocations and trade them for communication, recomputation, or transfer latency.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.