NVIDIA Expands Simulation Tools for Physical AI
NVIDIA outlines the growing role of simulation in training physical AI systems, with a focus on robotics and control policies.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
NVIDIA outlines the growing role of simulation in training physical AI systems, with a focus on robotics and control policies.
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max).
Kimi K3 scored 32.2% on ExploitBench, while leading U.S. models reached 76.2% in cyber exploit development, per a joint evaluation by the British AI Security Institute and the U.S. Center for AI Standards and Innovation.
Anthropic announced a $200 million fund to support research on AI's economic impacts, focusing on worker adaptation and policy interventions.
An AI system helped Pakistani judges clear 1,848 more cases per year per district, yielding a $38.50 return per dollar invested, according to a study.
Researchers suggest new AI evaluation metric to measure alignment with user intent, citing 80% of users struggle with current systems.
Amazon Nova 2 uses self-distilled reasoning to improve performance on math and coding tasks, recovering 70% of base model performance after SFT training.
Meta's Segment Anything Model 3 and DINOv3 enable real-time analysis of X-ray data, reducing segmentation tasks from weeks to 15 minutes at the Lawrence Berkeley National Laboratory.
Anthropic announced rare disease research grants offering up to $50,000 in Claude credits to scientists and biotechs. The initiative aims to accelerate discovery and improve treatment for rare genetic conditions.
LLMs like ChatGPT and Gemini showed stronger biases than humans in a simulated hiring experiment, scoring 65% higher on segregation scales.
A new benchmark shows AI models can be dangerously confident in wrong X-ray diagnoses, scoring 758 out of 2,000 points compared to human radiologists' 988.7.
AI text detectors like Pangram and GPTZero fail to catch 13% of style-mimicking AI texts, according to Epoch AI research.