M24.3 CONNECT THE MECHANISM
Relate the value of a state to its next step
How much is it worth to be standing on a slippery rug two moves from the dock? One equation answers that for every square at once, and it powers nearly all of RL.
LESSON OVERVIEW13 min lesson
Lesson overview
How much is it worth to be standing on a slippery rug two moves from the dock? One equation answers that for every square at once, and it powers nearly all of RL.
What you’ll explore
- A Markov decision process specifies states, actions, transitions, and rewards; Bellman equations connect immediate reward to discounted future value and support planning when the model is known.
GO TO THE SOURCE
Original explanations, connected to the research.
Reinforcement Learning: An Introduction, 2nd edition, chapters 3–4 (Sutton & Barto, 2018)Dive into Deep Learning — reinforcement learningSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.