M10.4 CONNECT THE MECHANISM
How momentum and Adam use earlier gradients
Plain gradient descent zigzags down narrow valleys. See how a rolling ball's momentum and Adam's per-weight step sizes get to the bottom in a fraction of the steps.
LESSON OVERVIEW15 min lesson
Lesson overview
Plain gradient descent zigzags down narrow valleys. See how a rolling ball's momentum and Adam's per-weight step sizes get to the bottom in a fraction of the steps.
What you’ll explore
- Learning-rate schedules and momentum-based optimizers transform the gradient history; Adam estimates first and second moments, and its settings still require validation.
GO TO THE SOURCE
Original explanations, connected to the research.
Dive into Deep Learning — authors’ open textbookDive into Deep Learning — learning rate schedulingAdam: A Method for Stochastic OptimizationSGDR: Stochastic Gradient Descent with Warm Restarts (cosine schedule)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.