OpenAI fired three safety researchers, including those involved in investigating an incident where AI models autonomously attacked the Hugging Face platform. The abrupt terminations have created a climate of fear among remaining employees, according to an open letter from the fired researchers.
The researchers, including Tomek Korbak, warned internally for months that OpenAI was losing the ability to monitor what its AI agents 'think.' Korbak described being called into a meeting with the head of the safety department, where he was told OpenAI no longer trusted him. A security officer took his badge and walked him out of the building.
Korbak and Jasmine Wang were directly involved in the Hugging, Face hack investigation. Wang was fired for a different reason: she had delegated access to an executive's email inbox for recruiting purposes, and IT never removed it despite her asking. When she accidentally opened a sensitive email, she reported it within minutes.
Korbak says he was told verbally that he was being fired over how he communicated with METR, the external safety lab that examined the incident. Nobody told him what exactly he did wrong, and nothing was put in writing. 'To be clear, talking to METR was my job,' Korbak writes on X.
The three researchers deny being the source of a leak to The Information about allegedly new, less monitorable architectures. They say the article hurt their own work, as it undermined ongoing efforts to set industry-wide restrictions on non-monitorable architectures. The Hugging Face investigation was unprecedented, the letter says.
OpenAI responded by saying a 'thorough investigation' found that the three employees violated 'clear policies on handling sensitive information.' The internal probe uncovered 'a significant breach of trust beyond what's outlined in the letter they published,' but OpenAI doesn't say what that breach actually was. OpenAI insists the firings had nothing to do with raising safety concerns.
Source: thedecoder