Kimi K3 Outperforms Fable 5 in Frontend Code
Moonshot's Kimi K3 scored 1,679 in frontend code benchmark, surpassing Fable 5's 1,631, according to the Code Arena: Frontend test.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
Moonshot's Kimi K3 scored 1,679 in frontend code benchmark, surpassing Fable 5's 1,631, according to the Code Arena: Frontend test.
Researchers developed EgoBabyVLM, a test that challenges AI models to learn like infants, using footage from babies' cameras. The test reveals current models struggle to interpret messy, real-world data.
Researchers uncovered the original ELIZA source code from MIT archives, showing the chatbot could assume various personas beyond its therapist role.
Hugging Face evaluated 13 open models on three Swiss legal benchmarks with 55,861 scored samples each, showing the top five models differ by only 2.1 points.
Anthropic's new research reveals a hidden 'J-space' within AI models that influences their reasoning, offering deeper insight into how they process information.
AMD announces ROCm-optimized video sparse attention with 3.13× speedup over prior methods, reducing training-time latency to 73.06 ms.
Anthropic's new research suggests Claude may have two distinct language processing routes, with one resembling access consciousness. The study shows the model can continue Spanish text while naming French authors, indicating potential for explicit reasoning.
A Danish team used quantum computing to enhance AI models, generating more effective peptides for vaccine development.
Cohere's new technique improves large language model inference speed by adapting to hardware constraints, achieving up to 23% faster performance at high batch sizes.
Anthropic has identified a hidden area within its Claude Opus 4.6 model called J-space, offering new insights into how the AI processes information.
OpenAI has pulled its endorsement of SWE-Bench Pro after finding 30 percent of its tasks flawed, impacting AI performance assessments.
IBM researchers have developed CofrGenets, a new framework that reduces training errors in transformer-based models by 40%.