Back to the lesson libraryMECHANISM · 14 MIN
M28.7 CONNECT THE MECHANISM

Test hypotheses about internal representations and computation

Researchers found a feature inside Claude that stood for the Golden Gate Bridge, then turned it up until the model claimed to be the bridge. Learn the tools that open a model up and test what's inside.

LESSON OVERVIEW14 min lesson

Lesson overview

Researchers found a feature inside Claude that stood for the Golden Gate Bridge, then turned it up until the model claimed to be the bridge. Learn the tools that open a model up and test what's inside.

What you’ll explore

  • Probes, activation interventions, circuit analyses, and sparse autoencoders provide different evidence about learned computation; distinguish decodability, causal influence, and complete mechanistic explanation.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.