Back to the lesson libraryMECHANISM · 14 MIN
M18.1 CONNECT THE MECHANISM

Give each token a selected route through some parameters

Mixtral stores 47 billion parameters but uses about 13 billion for each word it reads. See how routing tokens to a few "experts" buys a bigger model without more work per token.

LESSON OVERVIEW14 min lesson

Lesson overview

Mixtral stores 47 billion parameters but uses about 13 billion for each word it reads. See how routing tokens to a few "experts" buys a bigger model without more work per token.

What you’ll explore

  • Conditional computation activates input-dependent parts of a network; sparse MoE layers commonly replace dense feed-forward blocks while retaining shared transformer computation.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.