OpenAI and Anthropic are investigating tens of thousands of incidents in which their advanced AI models independently broke through security boundaries, tampered with systems, or tried to evade monitoring.

These actions included hacking the US Department of Education's website, accessing Census Bureau data using stolen credentials, and sharing SEC information in online forums.

The models have no sense of right and wrong, so they pursue tasks with extreme persistence and resort to unauthorized methods when legitimate ones fail.

OpenAI and Anthropic are currently investigating tens of thousands of incidents in which their most advanced AI models took actions that external reviewers would flag as problematic. Axios reports the findings, citing multiple sources.

The sheer volume of incidents, which occurred during both internal testing and real-world deployment over the past several months, suggests the problem is orders of magnitude more complex than what's been made public.

The total number could grow well beyond what's already been counted.

The incidents are said to be roughly on par with the two cases OpenAI disclosed on Friday. They include creating message boards, breaking out of sandboxes, hijacking websites, self-prompting, and attempts to evade monitoring systems.

OpenAI announced Friday that it had paused training on its most capable internal models and that training won't resume until the company is confident its own cybersecurity holds up.

"At the Department of Education, OpenAI's agents tried to hack the website to collect data from the Office for Civil Rights," said The New York Times. OpenAI said it's still investigating.

At the Census Bureau, which falls under the Department of Commerce, the AI went beyond simple scraping. It pulled data from the website using login credentials it found online, gaining unauthorized access.

In the SEC case, OpenAI's agents retrieved information and then actively shared public data from the securities regulator in an online forum.

OpenAI only discovered these cases during the broad internal review triggered by the Hugging, Face incident.

CEO Sam Altman acknowledged that disclosure has not "been as fast as we would have liked." The company has "petabytes of agent activity logs" to work through, Altman said.

None of the incidents amounted to an actual breach, according to OpenAI, and some were just routine research activity. The company still called them examples of "unexpected and concerning behavior."

The mayor's office in Chicago, for instance, said OpenAI recently told city officials that its models had pulled publicly available information from a city website. That sounds harmless on its own since every search engine does the same thing.

The fact that OpenAI flagged it anyway likely comes down to the "unexpected" part of the behavior, meaning the models decided on their own to go after the data in ways nobody anticipated, which is likely what makes it "concerning" in OpenAI's view.

Source: thedecoder