M27.4 CONNECT THE MECHANISM
Schedule many requests around shared hardware
One GPU, forty customers, replies of wildly different lengths. See how continuous batching and paged KV memory let one server hold six times as many conversations.
LESSON OVERVIEW14 min lesson
Lesson overview
One GPU, forty customers, replies of wildly different lengths. See how continuous batching and paged KV memory let one server hold six times as many conversations.
What you’ll explore
- Serving schedulers batch work, stream outputs, and manage reusable cache state; throughput gains must be evaluated alongside latency, fairness, memory fragmentation, and isolation.
GO TO THE SOURCE
Original explanations, connected to the research.
Orca: A Distributed Serving System for Transformer-Based Generative Models (Yu et al., OSDI 2022)Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., SOSP 2023), the vLLM paperSARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills (Agrawal et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.