A01 CONNECT THE MECHANISM

Follow a token through a mixture of experts

Some of the biggest language models use only a small slice of themselves for each word. Pick two "experts," blend their answers, and see what the trick costs.

LESSON OVERVIEW11 min lesson

Lesson overview

Some of the biggest language models use only a small slice of themselves for each word. Pick two "experts," blend their answers, and see what the trick costs.

What you’ll explore

  • In a sparse MoE layer, a router selects expert subnetworks and combines their outputs; active computation differs from total parameter storage.
GO TO THE SOURCE

Original explanations, connected to the research.

Mixtral of Experts (Jiang et al., 2024)DeepSeek-V3 Technical Report (DeepSeek-AI, 2024)
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.