World Models May Outpace Language Models in Scientific Discovery
Google DeepMind's Tom Zahavy argues language models lack the creative leap needed for scientific breakthroughs, citing Einstein's process as a benchmark.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
Google DeepMind's Tom Zahavy argues language models lack the creative leap needed for scientific breakthroughs, citing Einstein's process as a benchmark.
OpenAI reported that enabling two API settings increased GPT-5.6 Sol's score on the ARC-AGI-3 benchmark from 13.3% to 38.3%, surpassing human baselines.
Claude Opus 5 outperformed other models in a simulated vending machine business, earning $11,182 in its first year of operation.
LettucePrevent, a new tool from Hugging Face, cuts numeric hallucinations by more than 60% with minimal latency, according to a July 2026 blog post.
Siobahn Day Grady, a professor at North Carolina Central University, launched the first AI research institute at an HBCU in 2023.
Nearly a quarter of Deepmind's AlphaFold researchers have left the company, with several joining Anthropic. The shift marks a strategic pivot away from long-term scientific projects.
A self-taught AI engineer claims to have linked Linear A to Semitic languages using AI tools, compiling a 408-word lexicon based on a single hypothesis.
Anthropic's AI model found weaknesses in cryptographic algorithms, including a reduced version of AES, with no immediate impact on current systems.
Google's new study of 15 million AI interactions finds AI use remains shallow and collaborative, with less than 25% of work tasks significantly impacted.
Rohde & Schwarz highlights the limitations of static threat libraries against mode-agile emitters, which now dominate modern RF environments.
AWS announced task-aware knowledge compression (TAKC) to enhance enterprise AI workflows by pre-compressing knowledge bases into task-specific representations, reducing token count by 8x to 64x.
METR's new metric calculates when AI agents exceed human cost-effectiveness, estimating $2,500 per one-percent speedup in human effort.