M29.4 CONNECT THE MECHANISM
Distinguish manipulated inputs, training data, and model artifacts
A yellow sticker turns a stop sign into "speed limit," and a downloaded model file runs code the moment you open it. Find where each attack enters, and how to close the door.
LESSON OVERVIEW14 min lesson
Lesson overview
A yellow sticker turns a stop sign into "speed limit," and a downloaded model file runs code the moment you open it. Find where each attack enters, and how to close the door.
What you’ll explore
- Adversarial inputs, poisoning, backdoors, and compromised dependencies target different lifecycle stages; define attacker access and validate controls against the relevant threat model.
GO TO THE SOURCE
Original explanations, connected to the research.
Explaining and Harnessing Adversarial Examples (Goodfellow, Shlens & Szegedy, 2014)BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain (Gu, Dolan-Gavitt & Garg, 2017)Poisoning Web-Scale Training Datasets is Practical (Carlini et al., 2023)Obfuscated Gradients Give a False Sense of Security (Athalye, Carlini & Wagner, 2018)Compromised PyTorch-nightly dependency chain between December 25th and December 30th, 2022 (PyTorch)pickle — Python object serialization (warning on untrusted data)Safetensors documentation (Hugging Face)NIST AI Risk Management FrameworkSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.