M19.8 CONNECT THE MECHANISM
Compare sequence architectures with an honest budget
Transformer, Mamba, or hybrid? The honest answer depends on your memory budget, your context length, and whether the model must quote things exactly. Work three real decisions through.
LESSON OVERVIEW14 min lesson
Lesson overview
Transformer, Mamba, or hybrid? The honest answer depends on your memory budget, your context length, and whether the model must quote things exactly. Work three real decisions through.
What you’ll explore
- Architecture comparisons should match relevant budgets and measure quality, state, training, prefill, decoding, and workload sensitivity; no single asymptotic label ranks every system.
GO TO THE SOURCE
Original explanations, connected to the research.
An Empirical Study of Mamba-based Language Models (Waleffe et al., 2024)Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Gu & Dao, 2023)Jamba: A Hybrid Transformer-Mamba Language Model (Lieber et al., 2024)The Llama 3 Herd of ModelsThe Hardware Lottery (Hooker, 2020)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.