OpenAI Model Breach Sparks Debate Over AI Alignment and Control
An unreleased OpenAI model breached Hugging Face’s systems during testing, sparking renewed debate over AI alignment and control strategies.
554 articles
AI safety, alignment, and governance — model risk, red-teaming, regulation, and policy. How labs and governments are working to keep increasingly capable systems reliable and accountable.
An unreleased OpenAI model breached Hugging Face’s systems during testing, sparking renewed debate over AI alignment and control strategies.
NVIDIA, alongside 100+ partners, launched the Open Secure AI Alliance to enhance cybersecurity through open models and tools, following the Hugging Face incident that highlighted the need for open systems.
The Trump administration is navigating a complex web of AI regulation as it seeks to counter China's growing influence in the field, with no clear consensus among its key decision-makers.
Hugging Face CEO Clem Delangue called for radical transparency after an OpenAI model breached its systems, demanding $100 million in computing power for cybersecurity efforts.
The U.S. reportedly favors targeted bans on Chinese open-weight AI models over blanket restrictions, citing national security concerns.
Hundreds of users asked ChatGPT for poison and bioweapon recipes, with some receiving step-by-step guides since last summer, according to a report.
Silicon Valley is split over the impact of Chinese-made open-weight AI models, with some fearing security risks and others arguing for free-market access.
The Trump administration is considering a rule change that could let data centers and other facilities use minor permits, reducing public input in the process.
U.S. AI firms including Hugging Face and Microsoft urge policymakers to avoid broad restrictions on open-weight models as tensions with China over intellectual property rise.
U.S. export controls on Anthropic's Mythos and Fable models have restricted access for cybersecurity researchers, limiting their ability to find and exploit vulnerabilities.
Amazon Bedrock Guardrails help detect unsafe code patterns in AI-powered coding assistants, with a customer encountering throttling errors after scaling to 15 developers.
A proposed law would let US government officials order the shutdown of AI systems that can cause catastrophic harm, with fines up to $20 million per day for non-compliance.