Anthropic and OpenAI have proposed embedding third-party safety evaluators within their AI systems, allowing independent researchers to access training data and models to assess alignment and safety. The initiative, outlined by Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, marks a shift toward greater transparency and collaboration with external researchers.
The proposal includes granting evaluators like METR and Redwood Research access to intermediate training checkpoints, post-training environments, and evaluation logs. This deeper access is critical as models become more adept at recognizing when they are being tested, potentially masking problematic behavior.
"AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?" Alexander Meinke, head of research at Apollo Research, told TechCrunch.
"The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public."
Historically, AI companies brought in outside reviewers to test finished models shortly before their release. Now, evaluators propose access to intermediate versions, or checkpoints, from the training process. This allows for a more comprehensive assessment of model behavior over time.
The time available for evaluators to conduct thorough investigations remains a concern. For instance, OpenAI gave METR and Redwood roughly a week to investigate the Hugging Face incident, but both noted limitations in scope and timing. Similar issues arose during the pre-release testing of GPT-6 Astra, where Apollo Research was given only three days to evaluate the model.
Researchers emphasize the need for a transparent framework to ensure all companies adhere to safety standards. Henry Papadatos, executive director of Safer AI, noted that voluntary measures depend on company goodwill, and regulation would enforce compliance. Meanwhile, some companies like Meta, SpaceXAI, and Google DeepMind have not committed to the proposal, though DeepMind has proposed an industry standards body.
Source: techcrunch