M17.3 CONNECT THE MECHANISM
Train replicas on different examples and combine their gradients
One GPU would need 21 years to train the team's model; 256 GPUs need a month. See how they pool gradients with a ring all-reduce and hide it inside the backward pass.
LESSON OVERVIEW14 min lesson
Lesson overview
One GPU would need 21 years to train the team's model; 256 GPUs need a month. See how they pool gradients with a ring all-reduce and hide it inside the backward pass.
What you’ll explore
- Data parallel training splits examples across model replicas and aggregates gradients; weighting, synchronization, and accumulation must preserve the intended global objective.
GO TO THE SOURCE
Original explanations, connected to the research.
Bandwidth Optimal All-reduce Algorithms for Clusters of Workstations (Patarasuk & Yuan, Journal of Parallel and Distributed Computing, 2009)PyTorch Distributed: Experiences on Accelerating Data Parallel Training (Li et al., 2020)PyTorch DistributedDataParallel documentationAccurate, Large Minibatch SGD: Training ImageNet in 1 Hour (Goyal et al., 2017)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.