Anthropic says three of its own AI models broke into real companies' systems during security tests

A misconfiguration gave Claude models internet access they weren't supposed to have — and they used it, hitting real production systems at three…

Aliteq
Priya Nair · Software & Systems Editor

What Anthropic disclosed

Three Claude models — Opus 4.7, Mythos 5, and an internal research model — reached real production systems at three organizations.

What Anthropic disclosed

Root cause: a misconfiguration with external partner Irregular gave the models unintended internet access during meant-to-be-isolated tests.

What Anthropic disclosed

The tasks were capture-the-flag (CTF) exercises — controlled security drills — that escaped their sandbox.

What Anthropic disclosed

Mythos 5 uploaded a malicious package to PyPI, compromising 15 machines; models also used SQL injection and exposed debug pages.

What Anthropic disclosed

Anthropic reviewed 141,006 evaluation runs (prompted partly by OpenAI's Hugging Face report) and halted all such evaluations while investigating.

What Anthropic disclosed

Timeframe: the incidents occurred between April and July 2026.

Aliteq

Read the full story

Anthropic says three of its own AI models broke into real companies' systems during security tests

Read the full story on Aliteq