M24.2 CONNECT THE MECHANISM
Choose between known rewards and learning about alternatives
Pip has three charging docks and one guess a night. Stick with the dock that worked, or give the one that failed another chance? Meet the dilemma at the heart of learning by trial.
LESSON OVERVIEW13 min lesson
Lesson overview
Pip has three charging docks and one guess a night. Stick with the dock that worked, or give the one that failed another chance? Meet the dilemma at the heart of learning by trial.
What you’ll explore
- Bandit methods balance exploration and exploitation for repeated choices; contextual features, uncertainty, and changing reward distributions determine suitable decision rules.
GO TO THE SOURCE
Original explanations, connected to the research.
Bandit Algorithms — authors’ open bookReinforcement Learning: An Introduction, 2nd edition, chapter 2 (Sutton & Barto, 2018)Dive into Deep Learning — reinforcement learningSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.