Back to the lesson libraryMECHANISM · 14 MIN
M17.3 CONNECT THE MECHANISM

Train replicas on different examples and combine their gradients

One GPU would need 21 years to train the team's model; 256 GPUs need a month. See how they pool gradients with a ring all-reduce and hide it inside the backward pass.

LESSON OVERVIEW14 min lesson

Lesson overview

One GPU would need 21 years to train the team's model; 256 GPUs need a month. See how they pool gradients with a ring all-reduce and hide it inside the backward pass.

What you’ll explore

  • Data parallel training splits examples across model replicas and aggregates gradients; weighting, synchronization, and accumulation must preserve the intended global objective.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.