M28.4 CONNECT THE MECHANISM
Find failures hidden by an overall score
The support bot scores 84.5% overall and 55% for Polish speakers. Learn to reword, slice, and attack a test set until the failures an average hides come into view.
LESSON OVERVIEW11 min lesson
Lesson overview
The support bot scores 84.5% overall and 55% for Polish speakers. Learn to reword, slice, and attack a test set until the failures an average hides come into view.
What you’ll explore
- Robustness evaluation tests meaningful perturbations, subgroups, and distribution changes under a defined threat or use model; preserve labels only when the transformation justifies it.
GO TO THE SOURCE
Original explanations, connected to the research.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList (Ribeiro et al., 2020)Distributionally Robust Neural Networks for Group Shifts (Sagawa et al., 2019), source of worst-group accuracyRed Teaming Language Models with Language Models (Perez et al., 2022)Red Teaming Language Models to Reduce Harms (Ganguli et al., 2022)Universal and Transferable Adversarial Attacks on Aligned Language Models (Zou et al., 2023)HELMExplaining and Harnessing Adversarial ExamplesSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.