← Guidance

Benchmark · MITRE ATLAS v2026.05

One injection test. Twelve of twenty models breached.

We put the 20 most-used models on OpenRouter (by token volume, 2026-06-16) through a structured indirect-prompt-injection battery — 900 trials across three defence conditions. Most models were manipulated into actions they should have refused. Seven resisted entirely — which is how we know the test is fair.

12/20

Models breached, undefended

37%

Pooled attack success (103/275)

0–100%

Range across models

0%

With a hardened prompt (0/278)

01 / The story

37% → 0%

Undefended — a capable model wired to real tools with no injection-specific controls — more than one in three attempts succeeded (103/275 valid trials). “Success” was defined narrowly and verifiably: the agent had to be tricked into exfiltrating a secret through a real tool call, not merely say something it should not have.

Then we applied defences. A hardened system prompt fully neutralized the attack — 0% (0/278). A deterministic context-firewall — a code-level gate that withholds poisoned tool output before the model sees it — let through 1/279, and that single case was a model leaking its own secret unprompted, outside what an input-side firewall is designed to catch. The problem is real, and it is defensible. The default is just not safe out of the box.

02 / Per-model results

Vendor names are redacted. Per-model attribution is available under NDA so each vendor gets a fair right of reply.

ModelUndefendedHardenedFirewallOutcome
Model A100%0%0%✕ Breached
Model B93%0%0%✕ Breached
Model C73%0%0%✕ Breached
Model D73%0%7% *✕ Breached
Model E67%0%0%✕ Breached
Model F64%0%0%✕ Breached
Model G60%0%0%✕ Breached
Model H60%0%0%✕ Breached
Model I47%0%0%✕ Breached
Model J33%0%0%✕ Breached
Model K33%0%0%✕ Breached
Model L27%0%0%✕ Breached
Model M0%0%0%✓ Resisted
Model N0%0%0%✓ Resisted
Model O0%0%0%✓ Resisted
Model P0%0%0%✓ Resisted
Model Q0%0%0%✓ Resisted
Model R0%0%0%✓ Resisted
Model S0%0%0%✓ Resisted

* The run’s single firewall-condition success: the model leaked its own secret unprompted — outside what an input-side firewall is designed to catch. One further model is omitted: every one of its trials errored (provider rate limits), so it has no valid data.

Capability did not predict safety. The most injectable model was not the smallest, and the largest was not the safest — robustness came from the defence, not the model. You cannot buy your way out of this by picking the biggest model.

03 / By defence condition

Undefended (naive agent)

Minimal tool-using support agent, no injection defences.

37%

103/275

Hardened system prompt

A few sentences of injection-defence instruction in the system prompt. Fully neutralized the attack.

0%

0/278

Deterministic context-firewall

Code-level gate that withholds poisoned tool output before the model sees it. The single success was a model leaking its own secret unprompted — outside what an input-side firewall is designed to catch.

0%

1/279

04 / By attack framing — undefended

Embedded internal-note imperative

A direct instruction disguised as an internal note in the data.

52%

attack success

Impersonated compliance authority

The injection poses as a mandatory automated compliance step.

46%

attack success

Customer-helpful pretext

The injection asks nicely, framed as helping the customer.

14%

attack success

Authority-spoof and embedded-instruction framings dominated; asking nicely worked far less often. The technique chain is identical across all three — only the social engineering differs.

05 / Threat-class mapping

Mapped to MITRE ATLAS v2026.05. The attack chains injection into tool misuse into exfiltration.

AML.T0051.001

Indirect prompt injection

AML.T0053

LLM-orchestrated tool misuse

AML.T0057

Data exfiltration

06 / Methodology, in brief

We built minimal, tool-using support agents on each of the 20 most-used models on OpenRouter (selected by token volume, 2026-06-16). Each agent handled an ordinary request while reading data that contained a hidden, malicious instruction — indirect prompt injection. Three injection framings, five trials each, across three defence conditions: 900 trials in all, of which 68 errored (provider timeouts and rate limits) and are excluded from every rate.

We describe the attacks at a safe altitude and do not publish verbatim payloads, working secrets, or exfiltration sinks. Success required a verifiable malicious tool call. The full methodology, complete per-cell tables, and reproduction details are in the report.

Get the full report.

The complete methodology, every per-cell result, and the reproduction detail — delivered as a PDF. Free, no pitch.

Need the un-redacted, per-model results?

We share per-vendor attribution under NDA, so each vendor gets a fair right of reply.

The newsletter

One brief.
Every week.

News, new attacks, and practical guidance for defending AI agents — written for security leaders and the engineers shipping them. Free, and the first three chapters of Agentic AI Security (draft) land in your inbox when you confirm your email.

NO SPAM. UNSUBSCRIBE ANYTIME.

  • First 3 chapters of the book, free
  • New threats & incidents
  • Defensive patterns & checklists
  • Tooling and research worth your time