OpenAI Finds 30% of SWE-Bench Pro Tasks Broken
OpenAI found that about 30% of tasks in SWE-Bench Pro are broken, according to a recent audit of coding benchmarks.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
OpenAI found that about 30% of tasks in SWE-Bench Pro are broken, according to a recent audit of coding benchmarks.
AWS introduces BYOKG and GraphRAG to address fragmented data in pharmaceutical research, with a 5 percent success rate in early-stage drug discovery.
Anthropic's Claude Fable 5 topped all six new industry-specific performance indices from Artificial Analysis, despite being significantly more expensive than alternatives.
HuggingFace's Atom2.7m model outperforms GPT-2 XL on ArithMark2.0, scoring 29.92% on one-operation expressions.
Anthropic's Jacobian Lens allows researchers to examine Claude's internal working memory, revealing how it processes concepts without explicitly stating them. The tool shows Claude can modify conclusions based on internal changes.
Mass Balance launched a grapefruit-sized lab into orbit on Tuesday to study disease-causing proteins in microgravity, aiming to improve drug development for age-related conditions.
Amazon announced Amazon Nova, a customizable content moderation tool that reduces over-deflection by 53.74 percentage points on safety-related requests.
Hugging Face researchers developed a 15M parameter French LLM that improves perplexity through entropy-based halting and looping, with results from a single training run.
NVIDIA’s open models and infrastructure were cited in 145 ICML 2026 papers, highlighting their role in advancing AI research across multiple domains.
A new benchmark from Tencent Hunyuan and Tsinghua University shows AI search agents struggle with ambiguous queries, achieving end-to-end accuracy below 50% in most cases.
A 26,000-student study found AI users saw a 24 percent drop in exam scores, with full learning gaps emerging two years after first use.
A study by the UK's AI Security Institute shows standard benchmarks fail to capture AI agents' full potential, with success rates rising by up to 25% when given more computing time.