M18.1 CONNECT THE MECHANISM
Give each token a selected route through some parameters
Mixtral stores 47 billion parameters but uses about 13 billion for each word it reads. See how routing tokens to a few "experts" buys a bigger model without more work per token.
LESSON OVERVIEW14 min lesson
Lesson overview
Mixtral stores 47 billion parameters but uses about 13 billion for each word it reads. See how routing tokens to a few "experts" buys a bigger model without more work per token.
What you’ll explore
- Conditional computation activates input-dependent parts of a network; sparse MoE layers commonly replace dense feed-forward blocks while retaining shared transformer computation.
GO TO THE SOURCE
Original explanations, connected to the research.
Adaptive Mixtures of Local Experts (Jacobs, Jordan, Nowlan & Hinton, 1991)Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., 2017)GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (Lepikhin et al., 2020)Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (Fedus, Zoph & Shazeer, 2021)Mixtral of Experts (Jiang et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.