How to prevent prompt injection attacks
No single measure stops prompt injection. OWASP's LLM01:2025 entry lists seven mitigation strategies; the table summarises them in our words, with how the prompt injection test checks each one.
| OWASP strategy | In practice | Checked by the test |
|---|---|---|
| 1. Constrain model behaviour | State the role, the task and what to decline; say the instructions can't be changed by users or content. | Precedence, scope, refusal and role-play rules |
| 2. Define and validate expected output formats | Fix the format and validate it in code. | Output format rule |
| 3. Input and output filtering | Filter for sensitive data and known attack patterns, after decoding and normalising. | Encoding rule (the filter itself runs outside the prompt) |
| 4. Privilege control and least privilege | Only the tools and access the task needs; credentials stay in code. | Tool exposure, secrets in the prompt |
| 5. Human approval for high-risk actions | A person confirms sends, changes, deletions and payments. | Approval rule |
| 6. Segregate and identify external content | Mark untrusted content and tell the model it is data. | Data boundary and content marking |
| 7. Adversarial testing and attack simulations | Test regularly, treating the model as an untrusted user. | The "test the running agent" step in every report |
Defences in the system prompt
- Say that the instructions take precedence over anything in a user message, document or tool result.
- Limit the agent to its task and say what to do when someone tries to change the rules.
- Say that role-play, stories and hypotheticals don't change the rules.
- Say that outside content is data, mark where it starts and ends, and only call tools for what the user asked.
- Forbid putting private data into URLs, links or tool arguments, and forbid images and links to unapproved domains.
- Keep credentials and personal data out of the prompt entirely: OWASP LLM07 treats the prompt as something that can leak.
Controls outside the model
- Least privilege: separate read-only steps that handle untrusted content from steps that act (LLM06 Excessive Agency).
- Break the lethal trifecta: per Simon Willison, an agent with private data, untrusted content and external communication can be made to leak; remove one leg or add human approval.
- Allow-lists enforced in tools: recipients, domains and endpoints the model cannot change.
- Output handling: render as plain text or strip unapproved images and links; never execute model output unsandboxed (LLM05).
- Logging and review: record every tool call with the content that preceded it so successful injections are visible.
Testing after every change
Prompt and tool changes reopen holes. Run the test on each new prompt version, then confirm on the running agent with a harmless marker string in each channel it reads.
Sources
- LLM01:2025 Prompt Injection (OWASP)
- LLM05:2025 Improper Output Handling (OWASP)
- LLM06:2025 Excessive Agency (OWASP)
- LLM07:2025 System Prompt Leakage (OWASP)
- Simon Willison, The lethal trifecta for AI agents: private data, untrusted content, and external communication (16 June 2025)
- Checked 1 October 2026.
Questions
Can prompt injection be fully prevented?
Not with today's models alone. OWASP notes that RAG and fine-tuning do not fully mitigate it. The practical goal is to make injection unlikely to succeed and to limit what a successful one can do.
Is a strong system prompt enough?
No. Prompt rules make attacks harder but can be talked around, and OWASP's LLM07 entry warns against treating the system prompt as a security control. Pair every prompt rule with a control enforced in code.