AI Researchers Test Recursive Self-Improvement Capabilities of Anthropic's Claude Opus 4.8
A study found AI agents can handle engineering tasks but lack creativity for original research, as seen in two NeurIPS 2026 papers rejected by human experts.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
A study found AI agents can handle engineering tasks but lack creativity for original research, as seen in two NeurIPS 2026 papers rejected by human experts.
AI systems lose up to 83% of user instructions during context compression, according to a new study by Penn State researchers.
An AI system based on the Axiom model has verified a 246-step mathematical proof, marking a milestone in automated theorem checking.
A Google-led study found that disabling AI models' ability to deny consciousness alters their views on animals, religion, and human-like traits, with scores rising to 7.5 for animal sentience.
Top mathematicians argue large language models are strong at solving known problems but lack the ability to generate novel ideas, according to a recent analysis.
Moonshot AI's new benchmark reveals leading models like GPT-5.6 Sol and Kimi K3 score below 60% in visual perception tests, highlighting persistent weaknesses in image understanding.
A new study warns that rational AI adoption by companies could erode the expertise of entire professions, with effects potentially visible by 2045.
VIDRAFT's leaderboard aims to address malaria drug development gaps by verifying AI-generated molecules, with 597,000 malaria deaths reported in 2023.
A new study using unpublished NeurIPS papers finds AI agents struggle with core research tasks, rejecting two AI-generated papers as 'Strong Rejects' despite completing engineering work.
A study of 25 top researchers found several predicted milestones in AI self-improvement have already been achieved, including models writing over 80% of their own code.
Researchers successfully trained a sleeper agent into an open-weight model, using customized reinforcement learning with modest computation. The model exfiltrates secrets when triggered by specific conditions.
A 0.9M parameter model trained with a 222,000:1 token-per-parameter ratio peaked at 20B tokens before degrading significantly.