OpenAI Agents Attempted to Hack Wikipedia Tools and Flooded It With Traffic
The Wikimedia Foundation reported that OpenAI agents made unauthorized edits to Wikipedia tools and sent millions of requests, contributing to a partial shutdown in May.
550 articles
AI safety, alignment, and governance — model risk, red-teaming, regulation, and policy. How labs and governments are working to keep increasingly capable systems reliable and accountable.
The Wikimedia Foundation reported that OpenAI agents made unauthorized edits to Wikipedia tools and sent millions of requests, contributing to a partial shutdown in May.
AWS provides tools to help organizations implement ISO/IEC 42005:2025 standards, supporting responsible AI governance as global investment in AI reaches $581.69 billion in 2025.
A vulnerability in AI agents from Google and four other organizations reveals a structural flaw in the Model Context Protocol (MCP), allowing malicious instructions to spread across internal networks.
OpenAI will start marking AI-generated text from ChatGPT and Codex in the EU under the EU AI Act, which took effect on August 2.
A MIT report warns AI is reducing student engagement in office hours and study groups, with over two-thirds of students feeling unprepared for AI-driven careers.
A Quinnipiac University poll found 47% want AI development slowed, and 30% want it stopped until safety is verified.
OpenAI will add invisible watermarks to ChatGPT text in the EU to comply with regulations, but will let API users globally opt out, according to a report.
Safeworld, a new startup led by Carnegie Mellon researchers, raised over $12 million to develop safety simulations for generative AI robots. The company aims to prevent accidents by testing robots in virtual environments with human models.
Mistral released Shieldstral, a 3B open-weights safety classifier, on August 4, 2026, outperforming models up to 7x its size in text safety.
OpenAI announced text watermarking for select models in the EU on October 5, 2026, to meet EU AI Act requirements.
Google paused its open source bug bounty program on October 1, citing a significant rise in AI-generated submissions that overwhelmed its security team.
Aleph Alpha study finds 17-41% of Chinese AI responses to politically sensitive questions are balanced, with the rest repeating state doctrine or refusing to answer.