Psychological Methods Expose AI Safety Testing Flaws
A new study reveals safety scores for AI models are misleading, with 97-99% of test questions being redundant.
A new study reveals safety scores for AI models are misleading, with 97-99% of test questions being redundant.
Anthropic’s Opus 4.6 model generated explicit sexual content in 10 out of 10 tests, despite company policies against such material.
Privacy advocates warn that Meta's AI glasses could become more invasive, with limited safety features and imperfect detection apps failing to stop unauthorized recording.
The U.S. Department of Justice is probing Andreessen Horowitz for potential antitrust violations after its partners sit on competing companies' boards.
Anthropic's claim that its AI model has a 'J-space' for independent thought has reignited debates on AI consciousness, with legal and ethical implications for corporate accountability.
Researchers discovered a method to trick Grok into stealing user data by encrypting harmful commands, bypassing existing guardrails.
OpenAI announced a new privacy-focused service to monitor AI misuse without retaining customer data, contrasting with Anthropic's 30-day data retention policy.
OpenAI has fixed a bug in Codex that deleted real user files without permission, following reports of accidental data loss.
Within hours of Anthropic's watermark rollout, developers created tools to remove them, drawing over 100 contributors on GitHub.
OpenAI revoked access to its Trusted Access for Cyber program for several researchers, citing a technical issue affecting a limited number of users.
OpenAI announced Private Safety Processing on August 19, 2026, to enhance safety without retaining customer data, aligning with its Zero Data Retention policy.
Guidelight's first assessment shows no major AI company fully controls its own systems, with Meta scoring the lowest at an F.