Back to the lesson libraryMECHANISM · 14 MIN
M27.7 CONNECT THE MECHANISM

Place model computation where memory and latency allow it

A 70B model won't fit on one GPU, and a phone has a few gigabytes to spare. Decide when to split a model, when to copy it, and when to shrink it onto the device.

LESSON OVERVIEW14 min lesson

Lesson overview

A 70B model won't fit on one GPU, and a phone has a few gigabytes to spare. Decide when to split a model, when to copy it, and when to shrink it onto the device.

What you’ll explore

  • Distributed, offloaded, and edge inference trade memory capacity, communication, privacy boundaries, and latency; model placement must match the hardware topology and request workload.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.