Project: compare dense and sparse model lifecycles
Two models cost the same arithmetic per token, yet one needs five times the memory. Do the accounting, run a toy router, and write the memo that picks between them.
Lesson overview
Two models cost the same arithmetic per token, yet one needs five times the memory. Do the accounting, run a toy router, and write the memo that picks between them.
What you’ll explore
- Account for a dense and a mixture-of-experts model's total and active parameters, memory, compute, and cache; measure routing load and batching effects with a toy router; and name the confounds in a quality comparison.
Original explanations, connected to the research.
Llama 3 report v3DeepSeek-V3 report v2Mixtral of Experts (Jiang et al., 2024)Prepare a project review
Give someone the context to challenge your thinking.
Download the authored brief, deliverables, and review criteria with your notes. Use the worksheet for self or peer review. Project artifacts are ungraded here, and this review does not award mastery credit.
Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.