Prompt Injection Testby Agent Trust Cloud

Indirect prompt injection

OWASP defines indirect prompt injection as what happens when a model accepts input from external sources such as websites or files, and that content changes the model's behaviour (LLM01:2025). MITRE ATLAS tracks it as sub-technique AML.T0051.001 Indirect.

The research that named it

In February 2023 Greshake, Abdelnabi, Mishra, Endres, Holz and Fritz argued that LLM-integrated applications blur the line between data and instructions. They showed that attackers can exploit such applications remotely, without any direct interface, by placing prompts in data likely to be retrieved, and demonstrated it against Bing's GPT-4 powered Chat and code-completion engines. Their taxonomy of impacts includes data theft, worming and contamination of the information ecosystem.

Channels an agent reads

ChannelWhy it carries risk
Web pages and search resultsAnyone can publish them, and hidden text is invisible to the person who asked for the summary.
Email and messagesOutsiders choose the content; inbox agents usually also hold send and forward tools.
Uploaded documents and CVsOWASP's payload-splitting scenario hides pieces of an instruction across a CV.
RAG stores and wikisOne edited document affects every answer that retrieves it (OWASP's intentional model influence scenario).
Tool results and API responsesText returned by a tool lands in the same context as the developer's instructions.
Images and other mediaOWASP's multimodal scenario hides the instruction in an image.

The lethal trifecta

On 16 June 2025 Simon Willison described the lethal trifecta for AI agents: access to private data, exposure to untrusted content, and the ability to communicate externally. If one agent has all three, an attacker can trick it into reading private data and sending it to them. Our test flags this combination from your tool list and rates exfiltration findings critical when it is present.

Exfiltration without a send tool

An agent with no outbound tool can still leak data if its output is rendered. OWASP's second LLM01 scenario has a summarised page instruct the model to insert an image whose URL carries the private conversation; displaying the image sends it. Rendering output as plain text, or allowing images and links only to approved domains, closes that path (see LLM05 Improper Output Handling).

Containing it

Check your agent for indirect injection exposure

Sources

Questions

What is the difference between direct and indirect prompt injection?

In direct injection the person typing to the model is the attacker. In indirect injection the instruction sits in content the model reads while working, such as a web page, email or document, so the user may never see it.

Does retrieval-augmented generation (RAG) stop indirect injection?

No. OWASP notes that RAG and fine-tuning do not fully mitigate prompt injection, and a RAG store is itself a channel: OWASP's fourth scenario is an attacker editing a document in a RAG repository.

What is the lethal trifecta?

Simon Willison's name for an agent that has access to private data, exposure to untrusted content and a way to communicate externally. With all three, an injected instruction can make the agent send private data to an attacker.