M18.8 CONNECT THE MECHANISM
Serve sparse models under real traffic
Mixtral does a 13B model's arithmetic but needs a 47B model's memory, and at batch 16 it reads almost every expert on every step. Work out when MoE serving is cheap, and when it isn't.
LESSON OVERVIEW12 min lesson
Lesson overview
Mixtral does a 13B model's arithmetic but needs a 47B model's memory, and at batch 16 it reads almost every expert on every step. Work out when MoE serving is cheap, and when it isn't.
What you’ll explore
- MoE serving must manage resident capacity, token routing, load imbalance, expert batching, and transfers; measure prompt processing and decoding under realistic traffic and quality constraints.
GO TO THE SOURCE
Original explanations, connected to the research.
Mixtral of Experts (Jiang et al., 2024)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)Fast Inference of Mixture-of-Experts Language Models with Offloading (Eliseev & Mazur, 2023)Switch Transformers (Fedus, Zoph & Shazeer, 2021)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.