Back to the lesson libraryMECHANISM · 12 MIN
M18.7 CONNECT THE MECHANISM

Understand how sparse routes learn and cross device boundaries

A token's gradient reaches only the two experts it visited, plus the router through their gates. On eight GPUs, that token also crosses the network four times per layer. Count both.

LESSON OVERVIEW12 min lesson

Lesson overview

A token's gradient reaches only the two experts it visited, plus the router through their gates. On eight GPUs, that token also crosses the network four times per layer. Count both.

What you’ll explore

  • Selected expert outputs and routing weights provide gradient paths, while discrete selection requires a defined training treatment; expert-parallel execution exchanges activations and their backward gradients.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.