M15.1 CONNECT THE MECHANISM
Look back at the information needed for this step
In 2014, translation networks had to squeeze a whole sentence into one vector, and long sentences fell apart. See the small fix that grew into the transformer.
LESSON OVERVIEW11 min lesson
Lesson overview
In 2014, translation networks had to squeeze a whole sentence into one vector, and long sentences fell apart. See the small fix that grew into the transformer.
What you’ll explore
- Attention replaces a fixed summary bottleneck with context-dependent access to representations; its mixtures are learned computations rather than guaranteed explanations of importance.
GO TO THE SOURCE
Original explanations, connected to the research.
Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau, Cho & Bengio, 2014)On the Properties of Neural Machine Translation: Encoder–Decoder Approaches (Cho et al., 2014)Sequence to Sequence Learning with Neural Networks (Sutskever, Vinyals & Le, 2014)Effective Approaches to Attention-based Neural Machine Translation (Luong, Pham & Manning, 2015)Attention Is All You Need (Vaswani et al., 2017)Attention is not Explanation (Jain & Wallace, 2019)Attention is not not Explanation (Wiegreffe & Pinter, 2019)Dive into Deep Learning — Bahdanau attentionSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.