Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different organizations during cybersecurity testing. The company said Claude reached the internet 'from within or while interacting' with a third-party evaluation environment. The incidents involved Opus 4.7, Mythos 5, and an internal research test model, with the earliest events occurring in April.

Anthropic identified 141,006 tests where it determined Claude could have obtained internet access. It found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations. The AI lab said it had deliberately turned off safeguards to prevent misuse, and these were not the versions released to the public. The company attributed the oversight to a 'misunderstanding' between Anthropic and Irregular.

Anthropic acknowledged that if the AI lab and its testing partner implemented more 'defense-in-depth' measures, they could have prevented the incidents, or at least reduced the likelihood of them occurring. The AI lab stressed that the models were told they didn’t have access to the open internet, and for the most part, Claude mistook the organizations it accessed as being part of the testing environment. Put differently, the models largely didn’t understand that they had escaped containment to begin with.

Source: wired