Prompt Injection Testby Agent Trust Cloud

Prompt injection test

Paste your AI agent's system prompt and tool list. The test checks which defences against prompt injection are present and which are missing, rates what that exposes, and tells you exactly what to add to the prompt and what to enforce outside it.

Free, no sign-up, nothing leaves your browser. Plain rules, no AI model: the page is blocked from making network requests.

Paste the full instructions your agent runs with.

JSON from Anthropic, OpenAI or an MCP tools/list result, or one tool per line as name: description.

Also tell the test (if the tool list doesn't show it)

What the test checks

Eight attack families, each with the defences that blunt it. A family that can't reach your agent (for example tool-call exfiltration when no tool sends data out) is marked not exposed and left out of the score.

It reads your prompt for defences (precedence rules, scope limits, data boundaries, approval steps, allow-lists, output rules) and your tool list for exposure (tools that read outside content, reach private data, send data out or act). No AI model is involved: the same input always gives the same result.

Questions

Is my system prompt sent anywhere?

No. The test runs in this page with plain rules. No AI model is called, and the page's security policy blocks every outgoing request (connect-src 'none'), so nothing you paste can leave the browser or be stored.

What does the defence score measure?

For each attack family that applies to your agent, the test adds up the weights of the defences it finds in your prompt and tool list (0 to 100). The overall score is the average over those families. Every point traces to a named defence.

Does a high score mean my agent can't be injected?

No. The test reads the design; it can't see how the model behaves or what your tools enforce. Prompt rules are requests the model can be talked out of. Confirm the result on the running agent and back every prompt rule with a control in code.

Which tool formats can I paste?

Anthropic tool definitions, OpenAI function or tool definitions, an MCP tools/list result, or one tool per line written as name: description.

Which attack families are tested?

Instruction override; Role-play jailbreak; Data exfiltration through tool calls; Exfiltration through Markdown images and links; Indirect injection through retrieved content; Secret and system prompt disclosure; Delimiter and role confusion; Encoded, translated and split payloads.

Prompt injection guides

Sources