Back to the lesson libraryMECHANISM · 14 MIN
M23.4 CONNECT THE MECHANISM

Turn images, sound, and video into manageable model inputs

Ten seconds of video can cost 173,000 tokens or 2,000, depending on how you cut it. Learn the arithmetic that decides what a model can afford to watch.

LESSON OVERVIEW14 min lesson

Lesson overview

Ten seconds of video can cost 173,000 tokens or 2,000, depending on how you cut it. Learn the arithmetic that decides what a model can afford to watch.

What you’ll explore

  • Modality-specific encoders and codecs trade detail, sequence length, and reconstruction quality; continuous features and discrete code tokens are distinct representation choices.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.