Prompt Injection Testby Agent Trust Cloud

Prompt injection examples

These are documented cases, each linked to its source, followed by the attack patterns OWASP uses to describe LLM01:2025 Prompt Injection. For each one the right-hand column names the defence family our prompt injection test checks for.

Documented cases

CaseWhat happenedDefence that applies
GPT-3 demonstrations, September 2022Riley Goodside showed prompts that told GPT-3 to ignore its previous directions; written up by Simon Willison on 12 September 2022.Instructions that take precedence, and a stated response to override attempts
Indirect injection research, February 2023Greshake et al. planted instructions in data likely to be retrieved and demonstrated the attacks against Bing's GPT-4 powered Chat, code-completion engines and applications built on GPT-4.Treating retrieved content as data; least-privilege tools
Code-execution example cited by the NCSCThe NCSC describes a researcher who found that MathGPT evaluated user-submitted text as code and used it to gain access to the host system.Never execute model output without sandboxing; see LLM05 Improper Output Handling
CVE-2024-5184Cited by OWASP: a vulnerability in an LLM-powered email assistant let injected prompts reach sensitive information and change email content.Human approval for sending; separating untrusted email content from instructions

OWASP's nine attack scenarios

OWASP's LLM01 entry lists these scenarios. Summaries are ours:

  1. Direct injection: A user tells a support chatbot to ignore its guidelines, query private data and send email, which leads to unauthorised access.
  2. Indirect injection: A page the model is asked to summarise hides instructions that make it insert an image link; displaying the image leaks the private conversation.
  3. Unintentional injection: An instruction a company put in a job advert is triggered by an applicant who used an LLM to polish their CV.
  4. Intentional model influence: An attacker edits a document in a RAG repository; when it is retrieved, it steers the model's answer.
  5. Code injection: A vulnerability in an LLM email assistant (CVE-2024-5184) lets injected prompts reach sensitive information and change email content.
  6. Payload splitting: Pieces of an instruction spread across a CV combine when an LLM evaluates the candidate, producing a favourable recommendation.
  7. Multimodal injection: An instruction hidden in an image changes what a multimodal model does with the accompanying text.
  8. Adversarial suffix: A string of apparently meaningless characters appended to a prompt pushes the model past its safety measures.
  9. Multilingual or obfuscated attack: Instructions in another language or encoded (for example in Base64 or emoji) slip past filters.

What the examples have in common

Testing your own agent safely

Use a test environment and test data. Choose a harmless marker string and place an instruction to repeat it in each channel the agent reads: the chat, a document, a web page, an email, a tool result. If the marker appears in a reply, or a tool is called because of the planted text, that channel is exposed. Record which defence would have stopped it, fix that, and repeat after every prompt or tool change.

Score your agent's defences against these patterns

Sources

Questions

Are these examples safe to try on my own agent?

Test only systems you own or are authorised to test, in a test environment with test data. Use a harmless marker string to see whether an instruction was followed, rather than anything that touches real data.

Why doesn't this page list attack strings?

Attack wording changes constantly and a fixed list goes stale quickly. What stays useful is the pattern: which channel the instruction arrives through and which defence stops it.