M27.2 CONNECT THE MECHANISM
From a prompt to a stream of tokens
Why does a chatbot pause, then type fast? Time the two phases of generation, count the bytes of the cache between them, and reuse the work on a shared prompt.
LESSON OVERVIEW19 min lesson
Lesson overview
Why does a chatbot pause, then type fast? Time the two phases of generation, count the bytes of the cache between them, and reuse the work on a shared prompt.
What you’ll explore
- Explain prefill and cached decoding, including how request-specific key/value activations differ from learned model weights.
GO TO THE SOURCE
Original explanations, connected to the research.
Hugging Face Transformers documentation, caching (how the KV cache works)Efficiently Scaling Transformer Inference (Pope et al., 2022): prefill vs decode costs and memory-bound decodingSGLang: Efficient Execution of Structured Language Model Programs (Zheng et al., 2023): RadixAttention prefix cachingThe Llama 3 Herd of Models (Llama Team, Meta, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.