M25.7 CONNECT THE MECHANISM
Keep the authority to act outside retrieved content
A hotel's web page quietly tells Wayfarer to email your passport number to a stranger. Why might it obey, and what stops the email from ever leaving?
LESSON OVERVIEW14 min lesson
Lesson overview
A hotel's web page quietly tells Wayfarer to email your passport number to a stranger. Why might it obey, and what stops the email from ever leaving?
What you’ll explore
- Agent systems should bind actions to explicit permissions and task scope, treat external content as untrusted evidence, and enforce consequential boundaries in the execution layer.
GO TO THE SOURCE
Original explanations, connected to the research.
Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (Greshake et al., 2023)NIST AI Risk Management FrameworkThe lethal trifecta for AI agents: private data, untrusted content, and external communication (Simon Willison, June 2025)Defeating Prompt Injections by Design (Debenedetti et al., 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.