M24.8 CONNECT THE MECHANISM
Learn from recorded decisions without assuming safe exploration
You can't let a new robot learn by tumbling down the stairs. Learn from recorded driving instead, and discover why copying an expert drifts and how offline RL stays cautious.
LESSON OVERVIEW14 min lesson
Lesson overview
You can't let a new robot learn by tumbling down the stairs. Learn from recorded driving instead, and discover why copying an expert drifts and how offline RL stays cautious.
What you’ll explore
- Offline RL and imitation learn from logged behavior; gaps in that data and shifts at deployment limit how far you can safely improve on it.
GO TO THE SOURCE
Original explanations, connected to the research.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (Levine et al., 2020)A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger; Ross, Gordon & Bagnell, 2011)Conservative Q-Learning for Offline Reinforcement Learning (Kumar et al., 2020)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.