M24.7 CONNECT THE MECHANISM
Plan with learned dynamics when the full state is hidden
Pip wakes up after being carried to an unknown room. Learn how an agent can build its own map, keep track of where it probably is, and replan as the evidence comes in.
LESSON OVERVIEW13 min lesson
Lesson overview
Pip wakes up after being carried to an unknown room. Learn how an agent can build its own map, keep track of where it probably is, and replan as the evidence comes in.
What you’ll explore
- Model-based RL uses estimated dynamics for planning or imagined experience; partial observability requires history or belief state, and model uncertainty grows across hypothetical rollouts.
GO TO THE SOURCE
Original explanations, connected to the research.
Reinforcement Learning: An Introduction, 2nd edition, chapter 8 (Sutton & Barto, 2018)World Models (Ha & Schmidhuber, 2018)Mastering Diverse Domains through World Models (DreamerV3; Hafner et al., 2023)Mastering Atari, Go, chess and shogi by planning with a learned model (MuZero)Artificial Intelligence: A Modern Approach — authors’ materialsSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.