OpenAI announced it has determined that Astra, one of its upcoming models, may possess critical cybersecurity capabilities. The company said its internal evaluations over the past few days suggest significant advancements in agentic coding and cybersecurity. These findings, alongside expert assessments, led OpenAI to conclude it cannot rule out critical cyber capabilities under its Preparedness Framework. The company emphasized the importance of transparency with the public and cybersecurity communities about this potential shift in model capabilities.
According to OpenAI, the Preparedness Framework, first published in December 2023, guides the company in identifying and planning for emerging capabilities. The framework defines critical cybersecurity capabilities as those that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world systems without human intervention, or devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. Preliminary evaluations of Astra indicate strong enough performance that OpenAI cannot rule out Critical capability level at this time.
OpenAI stated that it has scaled up robustness testing of its safeguards and security controls to ensure they are appropriate for deploying these capabilities. The company also implemented stricter security controls for higher-capability models, including isolated testing environments, restricted network access, and enhanced model weight protections. Additionally, universal monitoring for risky actions and misalignment across all agentic applications of Astra has been introduced. OpenAI plans to work with government agencies and AI safety organizations to test the model's capabilities.
Source: openai