M23.7 CONNECT THE MECHANISM
Predict what may happen next, with or without actions
A model that's only 0.1 feet off per step can be nearly 3 feet off four seconds later. Meet world models, and the idea of predicting meaning instead of pixels.
LESSON OVERVIEW15 min lesson
Lesson overview
A model that's only 0.1 feet off per step can be nearly 3 feet off four seconds later. Meet world models, and the idea of predicting meaning instead of pixels.
What you’ll explore
- World models learn predictive representations or dynamics for observations and actions; latent prediction can support planning without requiring pixel generation or proving a complete causal model.
GO TO THE SOURCE
Original explanations, connected to the research.
World Models (Ha & Schmidhuber, 2018)A Path Towards Autonomous Machine Intelligence (LeCun, 2022)Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA; Assran et al., 2023)Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA; Bardes et al., 2024)Genie: Generative Interactive Environments (Bruce et al., 2024)V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (Assran et al., 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.