M27.6 CONNECT THE MECHANISM
Let a cheap draft propose tokens for a target model to verify
A small model guesses the next four tokens, and the big model checks all four in the time it takes to write one. The answers come out exactly as the big model would write them.
LESSON OVERVIEW13 min lesson
Lesson overview
A small model guesses the next four tokens, and the big model checks all four in the time it takes to write one. The answers come out exactly as the big model would write them.
What you’ll explore
- Speculative decoding proposes several tokens and verifies them with a target model; distribution-preserving variants require an exact acceptance/correction rule, while speed depends on acceptance and execution cost.
GO TO THE SOURCE
Original explanations, connected to the research.
Fast Inference from Transformers via Speculative Decoding (Leviathan, Kalman & Matias, 2023)Accelerating Large Language Model Decoding with Speculative Sampling (Chen et al., 2023)Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads (Cai et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.