M28.7 CONNECT THE MECHANISM
Test hypotheses about internal representations and computation
Researchers found a feature inside Claude that stood for the Golden Gate Bridge, then turned it up until the model claimed to be the bridge. Learn the tools that open a model up and test what's inside.
LESSON OVERVIEW14 min lesson
Lesson overview
Researchers found a feature inside Claude that stood for the Golden Gate Bridge, then turned it up until the model claimed to be the bridge. Learn the tools that open a model up and test what's inside.
What you’ll explore
- Probes, activation interventions, circuit analyses, and sparse autoencoders provide different evidence about learned computation; distinguish decodability, causal influence, and complete mechanistic explanation.
GO TO THE SOURCE
Original explanations, connected to the research.
Locating and Editing Factual Associations in GPT (Meng et al., 2022), which introduced ROME and causal tracingIn-context Learning and Induction Heads (Olsson et al., 2022)Toy Models of Superposition (Elhage et al., 2022)Towards Monosemanticity: Decomposing Language Models With Dictionary Learning (Bricken et al., 2023)Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet (Templeton et al., 2024)Designing and Interpreting Probes with Control Tasks (Hewitt & Liang, 2019)Towards Automated Circuit DiscoverySuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.