A01 CONNECT THE MECHANISM
Follow a token through a mixture of experts
Some of the biggest language models use only a small slice of themselves for each word. Pick two "experts," blend their answers, and see what the trick costs.
LESSON OVERVIEW11 min lesson
Lesson overview
Some of the biggest language models use only a small slice of themselves for each word. Pick two "experts," blend their answers, and see what the trick costs.
What you’ll explore
- In a sparse MoE layer, a router selects expert subnetworks and combines their outputs; active computation differs from total parameter storage.
GO TO THE SOURCE
Original explanations, connected to the research.
Mixtral of Experts (Jiang et al., 2024)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.