M23.4 CONNECT THE MECHANISM
Turn images, sound, and video into manageable model inputs
Ten seconds of video can cost 173,000 tokens or 2,000, depending on how you cut it. Learn the arithmetic that decides what a model can afford to watch.
LESSON OVERVIEW14 min lesson
Lesson overview
Ten seconds of video can cost 173,000 tokens or 2,000, depending on how you cut it. Learn the arithmetic that decides what a model can afford to watch.
What you’ll explore
- Modality-specific encoders and codecs trade detail, sequence length, and reconstruction quality; continuous features and discrete code tokens are distinct representation choices.
GO TO THE SOURCE
Original explanations, connected to the research.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT; Dosovitskiy et al., 2020)Neural Discrete Representation Learning (VQ-VAE; van den Oord, Vinyals & Kavukcuoglu, 2017)Zero-Shot Text-to-Image Generation (DALL·E; Ramesh et al., 2021)High Fidelity Neural Audio Compression (EnCodec; Défossez et al., 2022)Robust Speech Recognition via Large-Scale Weak Supervision (Whisper; Radford et al., 2022)ViViT: A Video Vision Transformer (Arnab et al., 2021)VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training (Tong et al., 2022)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.