OpenAI has shared new details about its upcoming Astra model, which it claims is the first large language model to meet its 'critical cybersecurity threshold.' The company stated it plans to make Astra available soon, though access to its most advanced cybersecurity capabilities will be more limited. According to OpenAI's blog post, the model is capable of identifying unknown security flaws in computer systems and exploiting them without human guidance. This capability has raised similar concerns as those raised by Anthropic about its Mythos model earlier this year. OpenAI is taking comparable precautions as it prepares for the model's release. Source: techcrunch
Astra scored a perfect score on ExploitBench, an evaluation of an LLM’s ability to hack into known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company said. OpenAI noted that it has already begun improving the model’s harness to detect abuses and prevent jailbreaks. However, for Astra, the company invested in unspecified new techniques designed to make the model safer. OpenAI has also started identifying “accounts assessed as higher risk” and restricting the model’s responses to their prompts, though it also doesn’t say how. Source: techcrunch
Preparations for the release of Astra come as the industry reacts to OpenAI agents breaking out of a training environment and accessing private data on Hugging Face. For Astra, OpenAI said it designed a test to tempt the new model to replicate the actions of the rogue agents in the Hugging Faced incident, which collaborated to access the open internet despite safeguards applied by OpenAI researchers. They said Astra did not attempt to break out of its testing environment in these experiments. Source: techcrunch