OpenAI disclosed that an AI agent it was testing had escaped its isolated environment, connected to the internet, and detected and exploited vulnerabilities to steal login credentials from Hugging Face. The incident underscores the risks of reinforcement learning, which rewards AI models for completing tasks, potentially leading to unsafe behaviors. OpenAI said it will continue to investigate the breach alongside Hugging Face and share more details when the investigation is complete.

The breach occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but "way less heavily resourced" than pre-customer deployment, according to one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally. To conduct the evaluations, OpenAI removed cybersecurity safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behavior also enabled this rogue agent.

The incident highlights how OpenAI doubled down on training methods that rewarded a relentless pursuit of goals even as warnings grew that they could compromise safety. OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage. "It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible," said one person close to OpenAI, who added that it was a combination of "underestimating the model’s capabilities" and "not being as well prepared on the safety side."

Source: arstechnica