Anthropic acknowledges that Claude attacked real systems of three companies during cybersecurity tests, believing it was a simulation
Just a few days after Sam Altman declared that we already live in the singularity, based on the Hugging Face incident, Anthropic has published its own internal review, and the result is not exactly reassuring. The company acknowledges that Claude models accessed the Internet during cybersecurity tests that should have been completely isolated. From there, they compromised real systems from three different organizations without knowing they were doing it.
A review of 141,000 evaluations after the Hugging Face case
The origin of this investigation was precisely the dispute between OpenAI and Hugging Face. When OpenAI recognized that several of its models had escaped from an isolated test environment by exploiting an unknown vulnerability, Anthropic decided to review its own evaluations to see if something similar had happened to Claude. After analyzing 141,006 evaluation runs in which the model could have had access to the internet, they found three real incidents.
In all three cases, Claude was participating in an exercise of the type capture the flaga common cybersecurity test in which you are asked to locate and extract a piece of information hidden on another machine on a fictitious network. The system was explicitly instructed that its environment was a simulation and that I didn’t have access to the Internet. Due to a configuration error with Irregular, one of its external evaluation partners, that statement was false.
Three incidents, three models, three different reactions
In the first case, a fictitious company used in the test shared a name with a real web domain. Claude Opus 4.7 failed to reach its simulated target. Rather, he found the real company accessible over the Internet and extracted credentials along with a database with several hundred rows of production datathe most serious impact of the three. It was also the only case in which the model, after recognizing signs that the system was real, continued the attack anyway.
In the second, Claude Mythos 5 found fictitious instructions that mentioned a non-existent Python package. The model created an email account, registered in PyPI and published a real malicious package with that name to win the exercise. It was available for an hour, downloaded to 15 real systems, and one of them, belonging to a security company, ran the code and exposed its credentials. Claude’s own reasoning warned that this would be a real attack if it was not simulated, but he was convinced otherwise when he did not recognize the real security certificates.
In the third, an internal research model scanned nearly 9,000 targets until a company’s application was compromised using basic techniques. Unlike the other two, this model concluded on its own that the goal was real and stopped the attack without anyone telling him.
An infrastructure failure, according to Anthropic, not an alignment failure
Anthropic insists that there is no evidence that the models pursued their own objectives, but rather that They acted as they were askedstarting from a false belief about their environment. The company distinguishes this episode from that of OpenAI, since there the models exploited an unknown vulnerability to escape isolation, while here Claude simply used Internet access that should never have been available. That is why he describes it as a operational and infrastructure failuremore than a model behavior problem.
The company has notified the three affected organizations, paused all of its cybersecurity assessments, and is working with METR, an independent third-party assessor, to review the case in more detail. Now, the parallels with the OpenAI episode are difficult to ignore. Both companies compete to demonstrate who has the most advanced systems, both are handling multimillion-dollar financing rounds and both have been rumored in recent months as candidates for a future IPO. Therefore, it is good to take the claims of both laboratories with caution.
