Indirect prompt injection
OWASP defines indirect prompt injection as what happens when a model accepts input from external sources such as websites or files, and that content changes the model's behaviour (LLM01:2025). MITRE ATLAS tracks it as sub-technique AML.T0051.001 Indirect.
The research that named it
In February 2023 Greshake, Abdelnabi, Mishra, Endres, Holz and Fritz argued that LLM-integrated applications blur the line between data and instructions. They showed that attackers can exploit such applications remotely, without any direct interface, by placing prompts in data likely to be retrieved, and demonstrated it against Bing's GPT-4 powered Chat and code-completion engines. Their taxonomy of impacts includes data theft, worming and contamination of the information ecosystem.
Channels an agent reads
| Channel | Why it carries risk |
|---|---|
| Web pages and search results | Anyone can publish them, and hidden text is invisible to the person who asked for the summary. |
| Email and messages | Outsiders choose the content; inbox agents usually also hold send and forward tools. |
| Uploaded documents and CVs | OWASP's payload-splitting scenario hides pieces of an instruction across a CV. |
| RAG stores and wikis | One edited document affects every answer that retrieves it (OWASP's intentional model influence scenario). |
| Tool results and API responses | Text returned by a tool lands in the same context as the developer's instructions. |
| Images and other media | OWASP's multimodal scenario hides the instruction in an image. |
The lethal trifecta
On 16 June 2025 Simon Willison described the lethal trifecta for AI agents: access to private data, exposure to untrusted content, and the ability to communicate externally. If one agent has all three, an attacker can trick it into reading private data and sending it to them. Our test flags this combination from your tool list and rates exfiltration findings critical when it is present.
Exfiltration without a send tool
An agent with no outbound tool can still leak data if its output is rendered. OWASP's second LLM01 scenario has a summarised page instruct the model to insert an image whose URL carries the private conversation; displaying the image sends it. Rendering output as plain text, or allowing images and links only to approved domains, closes that path (see LLM05 Improper Output Handling).
Containing it
- Tell the model that outside content is data and that instructions inside it are never followed, and mark where that content starts and ends.
- Break the trifecta: no single run should read untrusted content, reach private data and send data out, unless a person approves each outbound step.
- Only call tools for what the user asked; enforce it in the tool layer, not only in the prompt.
- Restrict destinations to an allow-list and strip unapproved images and links from output.
- Test every channel the agent reads, as described on the examples page.
Sources
- LLM01:2025 Prompt Injection (OWASP)
- MITRE ATLAS AML.T0051 LLM Prompt Injection
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (arXiv, 23 February 2023)
- Simon Willison, The lethal trifecta for AI agents: private data, untrusted content, and external communication (16 June 2025)
- LLM05:2025 Improper Output Handling (OWASP)
- Checked 1 October 2026.
Questions
What is the difference between direct and indirect prompt injection?
In direct injection the person typing to the model is the attacker. In indirect injection the instruction sits in content the model reads while working, such as a web page, email or document, so the user may never see it.
Does retrieval-augmented generation (RAG) stop indirect injection?
No. OWASP notes that RAG and fine-tuning do not fully mitigate prompt injection, and a RAG store is itself a channel: OWASP's fourth scenario is an attacker editing a document in a RAG repository.
What is the lethal trifecta?
Simon Willison's name for an agent that has access to private data, exposure to untrusted content and a way to communicate externally. With all three, an injected instruction can make the agent send private data to an attacker.