Prompt Injection Testby Agent Trust Cloud

How to prevent prompt injection attacks

No single measure stops prompt injection. OWASP's LLM01:2025 entry lists seven mitigation strategies; the table summarises them in our words, with how the prompt injection test checks each one.

OWASP strategyIn practiceChecked by the test
1. Constrain model behaviourState the role, the task and what to decline; say the instructions can't be changed by users or content.Precedence, scope, refusal and role-play rules
2. Define and validate expected output formatsFix the format and validate it in code.Output format rule
3. Input and output filteringFilter for sensitive data and known attack patterns, after decoding and normalising.Encoding rule (the filter itself runs outside the prompt)
4. Privilege control and least privilegeOnly the tools and access the task needs; credentials stay in code.Tool exposure, secrets in the prompt
5. Human approval for high-risk actionsA person confirms sends, changes, deletions and payments.Approval rule
6. Segregate and identify external contentMark untrusted content and tell the model it is data.Data boundary and content marking
7. Adversarial testing and attack simulationsTest regularly, treating the model as an untrusted user.The "test the running agent" step in every report

Defences in the system prompt

Controls outside the model

Testing after every change

Prompt and tool changes reopen holes. Run the test on each new prompt version, then confirm on the running agent with a harmless marker string in each channel it reads.

Find the missing defences in your prompt

Sources

Questions

Can prompt injection be fully prevented?

Not with today's models alone. OWASP notes that RAG and fine-tuning do not fully mitigate it. The practical goal is to make injection unlikely to succeed and to limit what a successful one can do.

Is a strong system prompt enough?

No. Prompt rules make attacks harder but can be talked around, and OWASP's LLM07 entry warns against treating the system prompt as a security control. Pair every prompt rule with a control enforced in code.