M19.6 CONNECT THE MECHANISM
Combine attention, recurrence, and convolution deliberately
Attention looks things up exactly but its cache keeps growing; Mamba's memory is fixed but blurry. Jamba keeps one attention layer in eight, and here is why that ratio works.
LESSON OVERVIEW14 min lesson
Lesson overview
Attention looks things up exactly but its cache keeps growing; Mamba's memory is fixed but blurry. Jamba keeps one attention layer in eight, and here is why that ratio works.
What you’ll explore
- Hybrid sequence models mix mechanisms with different memory and communication properties; the arrangement, shared dimensions, and evaluation determine whether the combination helps.
GO TO THE SOURCE
Original explanations, connected to the research.
Jamba: A Hybrid Transformer-Mamba Language Model (Lieber et al., 2024)An Empirical Study of Mamba-based Language Models (Waleffe et al., 2024)Repeat After Me: Transformers are Better than State Space Models at Copying (Jelassi et al., 2024)Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models (De et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.