M17.4 CONNECT THE MECHANISM
Split a model across devices in different ways
When one model won't fit on one GPU, you can slice each layer or stack the layers across GPUs. Each choice has a price, and you can calculate it.
LESSON OVERVIEW13 min lesson
Lesson overview
When one model won't fit on one GPU, you can slice each layer or stack the layers across GPUs. Each choice has a price, and you can calculate it.
What you’ll explore
- Tensor, pipeline, context, and expert parallelism divide different dimensions of model execution; each arrangement introduces specific communication and scheduling costs.
GO TO THE SOURCE
Original explanations, connected to the research.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (Shoeybi et al., 2019)GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism (Huang et al., 2018)Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (Narayanan et al., 2021)DeepSeek-V3 Technical ReportSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.