Anthropic Study Shows AI Can Build Exploits in Hours, Not Weeks
Anthropic's study reveals large language models can create exploits from security patches in hours, not weeks, with Mythos Preview achieving results in under six hours.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
Anthropic's study reveals large language models can create exploits from security patches in hours, not weeks, with Mythos Preview achieving results in under six hours.
Astrophysicist Chi-kwan Chan uses Codex to refine algorithms for simulating black hole plasma, improving computational efficiency in extreme physics research.
Research from Writer reveals that memory systems can degrade AI model performance by increasing the likelihood of incorrect answers, with models becoming more sycophantic as user preferences fill their context window.
A new energy-saving method for large language model training can reduce power consumption by up to 14 percent, according to IEEE Spectrum.
A 2023 study claiming 80% of U.S. workers face AI exposure has influenced global policy discussions, yet its limitations are growing as AI evolves.
A randomized trial in Sierra Leone found AI tools improved math scores by 1.2 to 1.7 years over eight weeks for 1,763 students.
Microsoft Research's Lens model, requiring one-fifth the compute of Z-Image, outperforms larger rivals on benchmarks.
Machine learning models now assist in weather forecasting, with the European Centre for Medium-Range Weather Forecasts (ECMWF) deploying its AIFS model in February 2025.
A new study shows larger OLMo models reliably learn rare tasks that smaller ones fail to grasp, even with extensive training.
Japanese startup Sakana AI has created a research lab focused on recursive self-improvement, aiming to develop AI systems that evolve and enhance their own capabilities.
Robot demonstrations often mislead about real capabilities, experts warn, as videos show humanoid robots performing tasks but struggle with generalization.
The Estonian Language Institute released a new benchmark ranking LLMs on their ability to resist Russian propaganda, with Anthropic's Claude models leading the pack.