Researchers from leading AI labs have raised concerns about the risks of automated AI research, with several milestones they predicted have already been reached. A study by IAPS fellow Severin Field interviewed 25 experts from organizations like OpenAI, Anthropic, and Google Deepmind, highlighting the growing urgency of the issue. Field’s analysis, published in The Attack Surface newsletter, reveals that the Task Horizon benchmark has shown significant progress, with the length of tasks AI agents can complete doubling every six months since 2019, and some analysts suggesting the pace has accelerated to every four months since 2024. The debate is not about whether self-improvement is occurring, but whether it is recursive and self-sustaining. Skeptics argue that breakthroughs in memory, creativity, or hypothesis testing are still needed, as paradigm-shifting ideas lack training data and validation. Field’s findings show that several milestones have already been reached, including OpenAI and Google Deepmind achieving gold-medal performance in the Math Olympiad, Sakana’s 'AI Scientist' producing a peer-reviewed paper, and Anthropic’s Claude writing over 80% of its own production code.
Only four of 20 respondents expect research-capable models to be released as public products, with half anticipating they will remain internal. Field warns of a potential 'incentive flip' where withholding models becomes more valuable than selling them, citing recent security incidents and government access restrictions as signs of this trend. He recommends congressional hearings, a government-run Task Horizon benchmark, and research on verifying international AI agreements to address these concerns.
Field notes that the debate has not yet reached Washington, while AI labs continue to advance. Recently, 1,224 employees from major AI companies, including OpenAI and Meta’s chief scientists, signed an open statement warning that their organizations may be nearing the point of automating AI research.
Source: thedecoder