Back to the lesson libraryMECHANISM · 13 MIN
M17.4 CONNECT THE MECHANISM

Split a model across devices in different ways

When one model won't fit on one GPU, you can slice each layer or stack the layers across GPUs. Each choice has a price, and you can calculate it.

LESSON OVERVIEW13 min lesson

Lesson overview

When one model won't fit on one GPU, you can slice each layer or stack the layers across GPUs. Each choice has a price, and you can calculate it.

What you’ll explore

  • Tensor, pipeline, context, and expert parallelism divide different dimensions of model execution; each arrangement introduces specific communication and scheduling costs.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.