Prompt Injection Testby Agent Trust Cloud

What is prompt injection?

Prompt injection is input that changes what a large language model does in a way its developer did not intend. OWASP puts it first in the Top 10 for LLM Applications 2025 as LLM01:2025 Prompt Injection, and splits it into two kinds: direct, where the user's own input changes the model's behaviour, and indirect, where content the model takes in from outside (a website, a file) does it. OWASP also notes that the effect can be unintentional as well as malicious.

Where the term comes from

In September 2022 Riley Goodside demonstrated GPT-3 prompts that told the model to ignore its previous directions. Simon Willison wrote the attack up on 12 September 2022 in Prompt injection attacks against GPT-3, and "prompt injection" became the common name. The comparison with SQL injection is about mixing trusted instructions and untrusted input in one channel. Unlike SQL, there is no equivalent of a parameterised query for natural language.

How the standards describe it

BodyWhere prompt injection appears
OWASPLLM01:2025, with direct and indirect types, nine attack scenarios and seven mitigation strategies
NISTNIST AI 100-2 E2025, the adversarial machine learning taxonomy (March 2025), which covers attacks on generative AI systems
MITRE ATLASAML.T0051 LLM Prompt Injection, with sub-techniques Direct (AML.T0051.000), Indirect (AML.T0051.001) and Triggered (AML.T0051.002)
UK NCSCThinking about the security of AI systems (30 August 2023): prompt injection is input designed to make the model behave in an unintended way, from offensive output to revealing confidential information

What it can lead to

The damage depends on what the model can reach. A model that only writes text can be made to write the wrong text. A model connected to tools can be made to use them: OWASP's first scenario is a support chatbot told to query private data and send email. That is why OWASP pairs prompt injection with Excessive Agency and Sensitive Information Disclosure, and why the size of the risk is set more by the agent's tools and permissions than by the wording of its prompt.

Why it is hard to fix

OWASP notes that retrieval-augmented generation and fine-tuning, while useful, do not fully mitigate prompt injection. Defences are layered: prompt rules that tell the model what to ignore, filtering, least-privilege tools, human approval for consequential actions, and regular adversarial testing. See how to prevent prompt injection.

Test your system prompt's defences

Sources

Questions

Is prompt injection a bug in one model?

No. It follows from how language models work: instructions and the content they process arrive as one stream of text, so any model that reads outside text can be steered by it.

Who first described prompt injection?

In September 2022 Riley Goodside showed GPT-3 prompts that told the model to ignore its previous directions, and Simon Willison wrote the attack up on 12 September 2022 as 'Prompt injection attacks against GPT-3'.