M18.6 CONNECT THE MECHANISM
Separate always-on experts from routed specialists
DeepSeek-V3 gives each token 1 always-on expert plus 8 picked from 256. See why many small experts beat a few big ones, and how a tiny bias keeps them all busy.
LESSON OVERVIEW14 min lesson
Lesson overview
DeepSeek-V3 gives each token 1 always-on expert plus 8 picked from 256. See why many small experts beat a few big ones, and how a tiny bias keeps them all busy.
What you’ll explore
- MoE designs can combine shared experts with routed experts and vary expert granularity; compare their aggregate widths, active selections, and communication instead of expert count alone.
GO TO THE SOURCE
Original explanations, connected to the research.
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models (Dai et al., 2024)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts (Wang et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.