Because an agent treats the text it ingests as instructions, an attacker can plant a directive in untrusted content, for example a lead's email that says to export contacts or change a record. The agent, holding real permissions, may act on it.
The defense is not to trust the model to ignore the injection; it is to constrain what any action can do: scoped permissions, a human approval gate on consequential writes, and an audit trail. Then even a successful injection inherits a narrow, reviewable verb.
How does prompt injection actually reach an agent?
Through the content the agent reads as part of doing its job. An email in the inbox it is summarising, a web page it is researching, a PDF attached to a ticket, a note someone typed into a CRM record. The attacker does not need access to your systems, only to something your agent will read, which is a far lower bar than compromising an account.
Why can prompt injection not simply be filtered out?
Because a language model does not have a reliable boundary between instructions and data. Everything arrives as text in the same context window, so an instruction embedded in retrieved content looks structurally like an instruction from you. Filters raise the cost of an attack and catch known patterns; they do not close the class of attack, which is why the mitigation has to sit at the action layer rather than the input layer.
What actually limits the damage from prompt injection?
Assume the agent will occasionally be successfully steered, and make that survivable. Treat every piece of retrieved content as untrusted. Scope the agent permissions so a hijacked run cannot reach anything expensive. Keep a human approval on consequential and customer-facing writes, so a manipulated agent still cannot send or delete on its own. The goal is bounding the blast radius, not achieving perfect input hygiene.
Related terms
From definition to a working system
Mindlyft is the approval and audit layer over your AI GTM agents, every action drafted, human-approved, reversible, and logged. Start with a free 30-minute GTM Engineering Review.
Get your free review