M18.4 CONNECT THE MECHANISM
Count total capacity and active work separately
Eight experts of 7 billion should make 56 billion, yet Mixtral 8x7B has 46.7 billion. Count parameters the way MoE papers do, and learn which number predicts memory and which predicts speed.
LESSON OVERVIEW12 min lesson
Lesson overview
Eight experts of 7 billion should make 56 billion, yet Mixtral 8x7B has 46.7 billion. Count parameters the way MoE papers do, and learn which number predicts memory and which predicts speed.
What you’ll explore
- Total parameters describe stored model capacity, while activated parameters approximate the subset used per token; neither alone determines memory, FLOPs, latency, or quality.
GO TO THE SOURCE
Original explanations, connected to the research.
Mixtral of Experts (Jiang et al., 2024)Switch Transformers (Fedus, Zoph & Shazeer, 2021)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)Scaling Laws for Neural Language Models (Kaplan et al., 2020): the 2N FLOPs-per-token estimateSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.