M23.6 CONNECT THE MECHANISM
Model motion and persistence across frames
Did the bus stop, or drive past? No single frame can say. See what it costs a model to watch time, and the trick that makes it about 8 to 28 times cheaper.
LESSON OVERVIEW14 min lesson
Lesson overview
Did the bus stop, or drive past? No single frame can say. See what it costs a model to watch time, and the trick that makes it about 8 to 28 times cheaper.
What you’ll explore
- Video models must represent spatial content and temporal relationships; understanding and generation require evaluation of motion, identity, causality, and consistency beyond individual frame quality.
GO TO THE SOURCE
Original explanations, connected to the research.
Is Space-Time Attention All You Need for Video Understanding? (TimeSformer; Bertasius, Wang & Torresani, 2021)ViViT: A Video Vision Transformer (Arnab et al., 2021)Video Diffusion Models (Ho et al., 2022)The "something something" video database for learning and evaluating visual common sense (Goyal et al., 2017)Video generation models as world simulators (OpenAI Sora technical report, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.