M28.6 CONNECT THE MECHANISM
Ask what an explanation method actually measures
The model sent this chat to a human. Which words made it do that? Split the credit with gradients and Shapley values, and learn the test that some famous saliency maps fail.
LESSON OVERVIEW16 min lesson
Lesson overview
The model sent this chat to a human. Which words made it do that? Split the credit with gradients and Shapley values, and learn the test that some famous saliency maps fail.
What you’ll explore
- Attributions and counterfactuals describe how a model responds to chosen baselines or changes; they don't by themselves reveal real-world causes or actions that are actually possible.
GO TO THE SOURCE
Original explanations, connected to the research.
Axiomatic Attribution for Deep Networks (Sundararajan et al., 2017), which introduced integrated gradientsA Unified Approach to Interpreting Model Predictions (Lundberg & Lee, 2017), which introduced SHAPSanity Checks for Saliency Maps (Adebayo et al., 2018)Not Just a Black Box: Learning Important Features Through Propagating Activation Differences (Shrikumar et al., 2016), on gradient × inputCounterfactual ExplanationsLIMESuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.