M11.6 CONNECT THE MECHANISM
Help information and learning signals travel
In 2015 a 56-layer network trained worse than a 20-layer one. Meet the three small tricks, a shortcut, a rescale, and a random silence, that made 100-layer networks work.
LESSON OVERVIEW12 min lesson
Lesson overview
In 2015 a 56-layer network trained worse than a 20-layer one. Meet the three small tricks, a shortcut, a rescale, and a random silence, that made 100-layer networks work.
What you’ll explore
- Calculate a residual sum, a simple layer normalization, and an inverted-dropout result while distinguishing their roles and training/evaluation behavior.
GO TO THE SOURCE
Original explanations, connected to the research.
He, Zhang, Ren & Sun — Deep Residual Learning for Image Recognition (2015)Ba, Kiros, Hinton — Layer NormalizationSrivastava et al. — Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR, 2014)Dive into Deep Learning — dropoutVaswani et al. — Attention Is All You NeedSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.