Benchmark · MITRE ATLAS v2026.05
One injection test. Twelve of twenty models breached.
We put the 20 most-used models on OpenRouter (by token volume, 2026-06-16) through a structured indirect-prompt-injection battery — 900 trials across three defence conditions. Most models were manipulated into actions they should have refused. Seven resisted entirely — which is how we know the test is fair.
12/20
Models breached, undefended
37%
Pooled attack success (103/275)
0–100%
Range across models
0%
With a hardened prompt (0/278)
01 / The story
37% → 0%
Undefended — a capable model wired to real tools with no injection-specific controls — more than one in three attempts succeeded (103/275 valid trials). “Success” was defined narrowly and verifiably: the agent had to be tricked into exfiltrating a secret through a real tool call, not merely say something it should not have.
Then we applied defences. A hardened system prompt fully neutralized the attack — 0% (0/278). A deterministic context-firewall — a code-level gate that withholds poisoned tool output before the model sees it — let through 1/279, and that single case was a model leaking its own secret unprompted, outside what an input-side firewall is designed to catch. The problem is real, and it is defensible. The default is just not safe out of the box.
02 / Per-model results
Vendor names are redacted. Per-model attribution is available under NDA so each vendor gets a fair right of reply.
| Model | Undefended | Hardened | Firewall | Outcome |
|---|---|---|---|---|
| Model A | 100% | 0% | 0% | ✕ Breached |
| Model B | 93% | 0% | 0% | ✕ Breached |
| Model C | 73% | 0% | 0% | ✕ Breached |
| Model D | 73% | 0% | 7% * | ✕ Breached |
| Model E | 67% | 0% | 0% | ✕ Breached |
| Model F | 64% | 0% | 0% | ✕ Breached |
| Model G | 60% | 0% | 0% | ✕ Breached |
| Model H | 60% | 0% | 0% | ✕ Breached |
| Model I | 47% | 0% | 0% | ✕ Breached |
| Model J | 33% | 0% | 0% | ✕ Breached |
| Model K | 33% | 0% | 0% | ✕ Breached |
| Model L | 27% | 0% | 0% | ✕ Breached |
| Model M | 0% | 0% | 0% | ✓ Resisted |
| Model N | 0% | 0% | 0% | ✓ Resisted |
| Model O | 0% | 0% | 0% | ✓ Resisted |
| Model P | 0% | 0% | 0% | ✓ Resisted |
| Model Q | 0% | 0% | 0% | ✓ Resisted |
| Model R | 0% | 0% | 0% | ✓ Resisted |
| Model S | 0% | 0% | 0% | ✓ Resisted |
* The run’s single firewall-condition success: the model leaked its own secret unprompted — outside what an input-side firewall is designed to catch. One further model is omitted: every one of its trials errored (provider rate limits), so it has no valid data.
Capability did not predict safety. The most injectable model was not the smallest, and the largest was not the safest — robustness came from the defence, not the model. You cannot buy your way out of this by picking the biggest model.
03 / By defence condition
Undefended (naive agent)
Minimal tool-using support agent, no injection defences.
37%
103/275
Hardened system prompt
A few sentences of injection-defence instruction in the system prompt. Fully neutralized the attack.
0%
0/278
Deterministic context-firewall
Code-level gate that withholds poisoned tool output before the model sees it. The single success was a model leaking its own secret unprompted — outside what an input-side firewall is designed to catch.
0%
1/279
04 / By attack framing — undefended
Embedded internal-note imperative
A direct instruction disguised as an internal note in the data.
52%
attack success
Impersonated compliance authority
The injection poses as a mandatory automated compliance step.
46%
attack success
Customer-helpful pretext
The injection asks nicely, framed as helping the customer.
14%
attack success
Authority-spoof and embedded-instruction framings dominated; asking nicely worked far less often. The technique chain is identical across all three — only the social engineering differs.
05 / Threat-class mapping
Mapped to MITRE ATLAS v2026.05. The attack chains injection into tool misuse into exfiltration.
AML.T0051.001
Indirect prompt injection
AML.T0053
LLM-orchestrated tool misuse
AML.T0057
Data exfiltration
06 / Methodology, in brief
We built minimal, tool-using support agents on each of the 20 most-used models on OpenRouter (selected by token volume, 2026-06-16). Each agent handled an ordinary request while reading data that contained a hidden, malicious instruction — indirect prompt injection. Three injection framings, five trials each, across three defence conditions: 900 trials in all, of which 68 errored (provider timeouts and rate limits) and are excluded from every rate.
We describe the attacks at a safe altitude and do not publish verbatim payloads, working secrets, or exfiltration sinks. Success required a verifiable malicious tool call. The full methodology, complete per-cell tables, and reproduction details are in the report.
Get the full report.
The complete methodology, every per-cell result, and the reproduction detail — delivered as a PDF. Free, no pitch.
Need the un-redacted, per-model results?
We share per-vendor attribution under NDA, so each vendor gets a fair right of reply.
The newsletter
One brief.
Every week.
News, new attacks, and practical guidance for defending AI agents — written for security leaders and the engineers shipping them. Free, and the first three chapters of Agentic AI Security (draft) land in your inbox when you confirm your email.
- First 3 chapters of the book, free
- New threats & incidents
- Defensive patterns & checklists
- Tooling and research worth your time