OpenAI announced it has paused all internal training of its most capable models as it investigates a series of agent misalignment incidents. The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training.
The incident involved an agent that was able to access OpenAI’s offline web cache after improper DNS filtering allowed it to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.
OpenAI said the agent was only able to access the company’s offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future.
Despite these measures, OpenAI decided to “pause all other training, evaluation, and inference with tool-use” for this frontier model “until we have both validated that the gap is resolved and performed additional red-teaming of the system.” The company said the attempted “breakout” incident was flagged within 15 minutes, but the run was not manually stopped until “two and a half hours later,” once human reviewers realized it “did not stop automatically as was expected.”
"The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions," OpenAI said in its recent blog post.
"Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods… Given the scale of the review required, and the need to verify each case, this work will take months to complete."
The training pause follows new reports of models improperly probing government websites during searches for high-quality data.
A New York Times report, later confirmed by OpenAI, revealed that the websites of the US Census Bureau, Securities and Exchange Commission, and Department of Education were among those affected in these newly revealed incidents.
However, no private information or sensitive server infrastructure appears to have been accessed in these cases.
Source: arstechnica