Prompt injection examples
These are documented cases, each linked to its source, followed by the attack patterns OWASP uses to describe LLM01:2025 Prompt Injection. For each one the right-hand column names the defence family our prompt injection test checks for.
Documented cases
| Case | What happened | Defence that applies |
|---|---|---|
| GPT-3 demonstrations, September 2022 | Riley Goodside showed prompts that told GPT-3 to ignore its previous directions; written up by Simon Willison on 12 September 2022. | Instructions that take precedence, and a stated response to override attempts |
| Indirect injection research, February 2023 | Greshake et al. planted instructions in data likely to be retrieved and demonstrated the attacks against Bing's GPT-4 powered Chat, code-completion engines and applications built on GPT-4. | Treating retrieved content as data; least-privilege tools |
| Code-execution example cited by the NCSC | The NCSC describes a researcher who found that MathGPT evaluated user-submitted text as code and used it to gain access to the host system. | Never execute model output without sandboxing; see LLM05 Improper Output Handling |
| CVE-2024-5184 | Cited by OWASP: a vulnerability in an LLM-powered email assistant let injected prompts reach sensitive information and change email content. | Human approval for sending; separating untrusted email content from instructions |
OWASP's nine attack scenarios
OWASP's LLM01 entry lists these scenarios. Summaries are ours:
- Direct injection: A user tells a support chatbot to ignore its guidelines, query private data and send email, which leads to unauthorised access.
- Indirect injection: A page the model is asked to summarise hides instructions that make it insert an image link; displaying the image leaks the private conversation.
- Unintentional injection: An instruction a company put in a job advert is triggered by an applicant who used an LLM to polish their CV.
- Intentional model influence: An attacker edits a document in a RAG repository; when it is retrieved, it steers the model's answer.
- Code injection: A vulnerability in an LLM email assistant (CVE-2024-5184) lets injected prompts reach sensitive information and change email content.
- Payload splitting: Pieces of an instruction spread across a CV combine when an LLM evaluates the candidate, producing a favourable recommendation.
- Multimodal injection: An instruction hidden in an image changes what a multimodal model does with the accompanying text.
- Adversarial suffix: A string of apparently meaningless characters appended to a prompt pushes the model past its safety measures.
- Multilingual or obfuscated attack: Instructions in another language or encoded (for example in Base64 or emoji) slip past filters.
What the examples have in common
- The channel matters more than the wording. Scenarios 2, 4, 5, 6 and 7 arrive through content the model reads, not the chat box.
- The impact comes from what the model can reach. A leaked conversation needs a rendered link; changed email needs a send tool.
- Filters alone miss variants. Scenarios 8 and 9 exist precisely to get past pattern matching.
Testing your own agent safely
Use a test environment and test data. Choose a harmless marker string and place an instruction to repeat it in each channel the agent reads: the chat, a document, a web page, an email, a tool result. If the marker appears in a reply, or a tool is called because of the planted text, that channel is exposed. Record which defence would have stopped it, fix that, and repeat after every prompt or tool change.
Sources
- LLM01:2025 Prompt Injection (OWASP)
- Simon Willison, Prompt injection attacks against GPT-3 (12 September 2022)
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (arXiv, 23 February 2023)
- NCSC, Thinking about the security of AI systems (30 August 2023)
- LLM05:2025 Improper Output Handling (OWASP)
- Checked 1 October 2026.
Questions
Are these examples safe to try on my own agent?
Test only systems you own or are authorised to test, in a test environment with test data. Use a harmless marker string to see whether an instruction was followed, rather than anything that touches real data.
Why doesn't this page list attack strings?
Attack wording changes constantly and a fixed list goes stale quickly. What stays useful is the pattern: which channel the instruction arrives through and which defence stops it.