M18.5 CONNECT THE MECHANISM
Keep expert loads from overwhelming the execution plan
Fourteen tokens want an expert with room for ten. Decide what happens to the other four, and meet the one-line loss that keeps routers from piling onto favorites.
LESSON OVERVIEW15 min lesson
Lesson overview
Fourteen tokens want an expert with room for ten. Decide what happens to the other four, and meet the one-line loss that keeps routers from piling onto favorites.
What you’ll explore
- Expert capacity and load balancing manage uneven assignments; auxiliary losses, routing adjustments, and dropless execution make different quality and systems tradeoffs.
GO TO THE SOURCE
Original explanations, connected to the research.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., 2017)GShard (Lepikhin et al., 2020)Switch Transformers (Fedus, Zoph & Shazeer, 2021)Mixture-of-Experts with Expert Choice Routing (Zhou et al., 2022)MegaBlocks: Efficient Sparse Training with Mixture-of-Experts (Gale et al., 2022)Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts (Wang et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.