Back to the lesson libraryCOMPARISON · 13 MIN
M35.3 CONNECT THE MECHANISM

Compare Qwen3 dense and MoE models with reasoning modes

Add "/no_think" to a question and Qwen3 answers at once; leave it off and it reasons at length first, with the same weights. See how that switch was trained, and why the huge MoE caches less than the dense 32B.

LESSON OVERVIEW13 min lesson

Lesson overview

Add "/no_think" to a question and Qwen3 answers at once; leave it off and it reasons at length first, with the same weights. See how that switch was trained, and why the huge MoE caches less than the dense 32B.

What you’ll explore

  • The original Qwen3 release connects dense and MoE architecture choices with staged pretraining, reasoning-oriented adaptation, mode fusion, and distillation; compare exact checkpoints and mode-specific inference budgets.
GO TO THE SOURCE

Original explanations, connected to the research.

Qwen3 report v1Qwen3 official repository
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.