M27.8 CONNECT THE MECHANISM
Choose a deployment using quality and workload evidence
The average wait looks great, yet one customer in a hundred waits over three seconds. Measure what users actually feel, and work out what a million tokens really costs you.
LESSON OVERVIEW13 min lesson
Lesson overview
The average wait looks great, yet one customer in a hundred waits over three seconds. Measure what users actually feel, and work out what a million tokens really costs you.
What you’ll explore
- Measure serving by latency percentiles, throughput, memory, reliability, and cost under a stated workload, and compare systems at the quality level users actually need.
GO TO THE SOURCE
Original explanations, connected to the research.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving (Zhong et al., 2024)MLPerf Inference Benchmark (Reddi et al., 2019)Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., 2023)The Tail at Scale (Dean & Barroso, Communications of the ACM, 2013)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.