M24.4 CONNECT THE MECHANISM
Update values from sampled transitions
Pip doesn't know how slippery the rug is. It can still learn the value of every square, one step at a time, by learning from its own surprise.
LESSON OVERVIEW13 min lesson
Lesson overview
Pip doesn't know how slippery the rug is. It can still learn the value of every square, one step at a time, by learning from its own surprise.
What you’ll explore
- Monte Carlo methods learn from completed returns, temporal-difference methods bootstrap from estimates, and Q-learning uses an off-policy greedy next-action target under defined convergence assumptions.
GO TO THE SOURCE
Original explanations, connected to the research.
Reinforcement Learning: An Introduction, 2nd edition, chapter 6 (Sutton & Barto, 2018)Learning to predict by the methods of temporal differences (Sutton, 1988)Dive into Deep Learning — reinforcement learningSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.