M27.7 CONNECT THE MECHANISM
Place model computation where memory and latency allow it
A 70B model won't fit on one GPU, and a phone has a few gigabytes to spare. Decide when to split a model, when to copy it, and when to shrink it onto the device.
LESSON OVERVIEW14 min lesson
Lesson overview
A 70B model won't fit on one GPU, and a phone has a few gigabytes to spare. Decide when to split a model, when to copy it, and when to shrink it onto the device.
What you’ll explore
- Distributed, offloaded, and edge inference trade memory capacity, communication, privacy boundaries, and latency; model placement must match the hardware topology and request workload.
GO TO THE SOURCE
Original explanations, connected to the research.
The Llama 3 Herd of Models (Llama Team, Meta, 2024), section on inferenceMegatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (Shoeybi et al., 2019)Efficiently Scaling Transformer Inference (Pope et al., 2022)FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU (Sheng et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.