Back to the lesson libraryMECHANISM · 12 MIN
M11.6 CONNECT THE MECHANISM

Help information and learning signals travel

In 2015 a 56-layer network trained worse than a 20-layer one. Meet the three small tricks, a shortcut, a rescale, and a random silence, that made 100-layer networks work.

LESSON OVERVIEW12 min lesson

Lesson overview

In 2015 a 56-layer network trained worse than a 20-layer one. Meet the three small tricks, a shortcut, a rescale, and a random silence, that made 100-layer networks work.

What you’ll explore

  • Calculate a residual sum, a simple layer normalization, and an inverted-dropout result while distinguishing their roles and training/evaluation behavior.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.