Back to the lesson libraryMECHANISM · 12 MIN
M15.6 CONNECT THE MECHANISM

Separate mixing across tokens from transforming one token

Every transformer layer repeats one small recipe: normalize, attend, add; normalize, MLP, add. Trace it with real numbers and count the 7 million weights in each GPT-2 layer.

LESSON OVERVIEW12 min lesson

Lesson overview

Every transformer layer repeats one small recipe: normalize, attend, add; normalize, MLP, add. Trace it with real numbers and count the 7 million weights in each GPT-2 layer.

What you’ll explore

  • Transformer blocks combine attention, position-wise feed-forward computation, residual additions, and normalization; gated activations and placement conventions change the exact computation.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.