OpenAI released its principles for third-party AI safety assessments on September 22, 2026, stating that these assessments are critical for balancing the responsibility of training, evaluating, and deploying models safely. It is the company's first major update on AI safety practices since its initial collaboration with third-party assessors in 2023.

OpenAI reported that its principles focus on four priority areas for deeper assessment, including the evaluation of safety cases, critical safeguards, and capability evaluations covering key risk categories. That compares with its earlier work on the Preparedness Framework, which included evaluations for cybersecurity and biological risks.

The principles are built on OpenAI's ongoing efforts to ensure transparency and accountability in AI development and deployment. Availability of these principles begins with public release, initially for independent assessment organizations in the private and non-profit sector.

"Third party assessments are a critical part of balancing that responsibility, expanding opportunities for input on AI safety, keeping the world informed, and keeping labs accountable to clear and independently supported safety claims," said Mira Murati, Chief Executive Officer of OpenAI. These assessments complement the company's work with governments on testing and evaluation, where distinct roles and responsibilities may call for different approaches.

The announcement follows OpenAI's public release of its Preparedness Framework in 2023. OpenAI itself frames the significance of these principles as a step toward ensuring that safety and security practices meet international standards.

OpenAI did not say how these assessments will be implemented, and it raised the open question of whether the current safeguards are sufficient to address all potential risks. The company said it will continue to work with independent assessors to refine and improve its safety practices.

Source: openai