Back to the lesson libraryMECHANISM · 14 MIN
M18.6 CONNECT THE MECHANISM

Separate always-on experts from routed specialists

DeepSeek-V3 gives each token 1 always-on expert plus 8 picked from 256. See why many small experts beat a few big ones, and how a tiny bias keeps them all busy.

LESSON OVERVIEW14 min lesson

Lesson overview

DeepSeek-V3 gives each token 1 always-on expert plus 8 picked from 256. See why many small experts beat a few big ones, and how a tiny bias keeps them all busy.

What you’ll explore

  • MoE designs can combine shared experts with routed experts and vary expert granularity; compare their aggregate widths, active selections, and communication instead of expert count alone.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.