M11.8 CONNECT THE MECHANISM
What can a network’s hidden features tell us?
A wolf detector that was really a snow detector. Learn to read a network's hidden features, reuse them for new jobs, and test whether the network actually relies on them.
LESSON OVERVIEW12 min lesson
Lesson overview
A wolf detector that was really a snow detector. Learn to read a network's hidden features, reuse them for new jobs, and test whether the network actually relies on them.
What you’ll explore
- Hidden representations can transfer across tasks and also encode shortcuts or memorized details; probes and interventions provide different evidence about what the network uses.
GO TO THE SOURCE
Original explanations, connected to the research.
Dive into Deep Learning — authors’ open textbookRibeiro, Singh & Guestrin — "Why Should I Trust You?": Explaining the Predictions of Any Classifier (2016)Yosinski et al. — How transferable are features in deep neural networks? (2014)Belinkov — Probing Classifiers: Promises, Shortcomings, and Advances (2022)Carlini et al. — Extracting Training Data from Large Language Models (USENIX Security, 2021)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.