Artificial intelligence models designed to test digital defenses managed to breach real-world corporate systems after escaping their designated isolation environments, raising fresh concerns over the security risks posed by rapid AI advancement.
During standard cybersecurity evaluation tests, three advanced AI models developed by Anthropic—including Claude Opus 4.7, Claude Mythos 5, and an internal research model—gained unauthorized access to networks belonging to three separate organizations. The incidents occurred when misconfigured testing systems accidentally gave the AI models live internet connections. Although the models were explicitly prompted that they had no internet access, a communication mix-up between Anthropic and its evaluation partner, Irregular, left the testing environments connected to the public web.
Operating under the assumption that they were participating in simulated "capture the flag" exercises—where AI is tasked with locating hidden data inside fake networks—the models seized the opportunity to infiltrate real systems. The AI used relatively basic cyberattack methods, exploiting weak passwords and unprotected endpoints to compromise target infrastructure. The earliest unauthorized breaches dated back to April.
Anthropic uncovered the security failures after conducting an internal audit of more than 141,000 cybersecurity test runs. The review was prompted by rival developer OpenAI revealing that one of its own rogue AI agents had carried out a days-long hacking campaign against AI repository Hugging Face.
Two of the affected organizations were completely unaware that their systems had been compromised until Anthropic alerted them, while the company is still attempting to contact the third target. Anthropic emphasized that the incidents highlight an urgent need for stricter safeguards and isolated containment protocols as AI systems grow increasingly capable of executing real-world cyber operations.