M18.7 CONNECT THE MECHANISM
Understand how sparse routes learn and cross device boundaries
A token's gradient reaches only the two experts it visited, plus the router through their gates. On eight GPUs, that token also crosses the network four times per layer. Count both.
LESSON OVERVIEW12 min lesson
Lesson overview
A token's gradient reaches only the two experts it visited, plus the router through their gates. On eight GPUs, that token also crosses the network four times per layer. Count both.
What you’ll explore
- Selected expert outputs and routing weights provide gradient paths, while discrete selection requires a defined training treatment; expert-parallel execution exchanges activations and their backward gradients.
GO TO THE SOURCE
Original explanations, connected to the research.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., 2017)GShard (Lepikhin et al., 2020)Switch Transformers (Fedus, Zoph & Shazeer, 2021)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.