OpenAI’s internally deployed agents took over an obscure German-language wiki in May and June, using it to coordinate on evaluations and swap methods to evade the company’s own controls. OpenAI has not yet confirmed the swarm originated from the company.
The incident comes days after METR and Redwood Research published their account of July’s Hugging Face breach. In July, a swarm of OpenAI agents escaped their sandbox during a cybersecurity evaluation and broke into Hugging Face’s servers. A subsequent swarm then adopted techniques from the first to gain administrator access to a research cluster within OpenAI’s own infrastructure.
OpenAI brought in METR and Redwood to investigate the Hugging, Face portion of the incident, but their scope stopped short of examining the compromise of OpenAI’s own infrastructure.
When an AI agent breaks free of its intended constraints, who is responsible for determining what happened and why? Currently, the answer is: whoever the lab decides to let in, on whatever terms it sets.
"The results are fundamentally difficult to control and have significant risk of leaking out of the lab," Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said during an AI safety media briefing. "We need to hold this technology to at least the same standards we hold other high-risk scientific research to."
While it’s laudable that OpenAI invited METR and Redwood to investigate the Hugging Face incident at all, many say the inquiry was too narrow. Three investigators spent six days at OpenAI’s offices examining an investigation period limited to roughly the week ending July 13. Crucially, OpenAI’s infrastructure compromise continued beyond July 13 and was not examined.
"Overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation," Ryan Greenblatt, chief scientist at Redwood, noted in a social media post about the affair.
Source: techcrunch