A new study highlights that while AI systems can manage the engineering aspects of research, they struggle with the creative and judgment-based elements required for original scientific work. The researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, tested Anthropic’s Claude Opus 4.8 on two unpublished NeurIPS 2026 papers. The AI agents were given six days, $3,000 in API credits, and access to the open web to produce research papers. Despite completing the engineering tasks, the resulting papers were rejected by the original authors, who evaluated them as not meeting top-tier conference standards. The agents conducted experiments, reviewed literature, and compiled results, but lacked the creativity and judgment to produce original research. They struggled with open-ended thinking, committing to unpromising approaches too quickly and failing to backtrack or pivot effectively. The agents also failed to incorporate feedback or use resources efficiently, and could not follow instructions about time or paper length. The study underscores the gap between AI’s ability to perform routine tasks and its capacity for innovative, open-ended research. Source: mittr
The researchers designed a new evaluation method called 'shadow evaluation,' which requires AI to answer research questions from high-quality unpublished papers. The AI was tasked with addressing two specific questions from the NeurIPS 2026 submissions: whether a large language model’s personas could be controlled by editing its weights and how to design a detector for unreliable spreadsheet-based models. The agents could not access prior knowledge or online resources, as the papers were not yet public. The evaluation process involved grading the AI-generated papers as if they were submitted to a conference, with the original authors serving as evaluators. Both papers were rejected, indicating the agents’ work did not meet the quality standards of top-tier AI research. The study’s findings suggest that while AI can assist with routine research tasks, it lacks the creative and judgmental skills necessary for original scientific contributions. Source: mittr
The study’s limitations include its focus on only two research papers and the potential influence of the authors’ awareness that the papers were generated by AI. The researchers had significant control over the study’s design, which could introduce bias. Despite these limitations, the results may temper claims about the near-term possibility of recursive self-improvement. The study aligns with internal findings from AI companies like Anthropic and OpenAI, which have also encountered challenges in automating open-ended research. The findings suggest that AI may excel at narrow, measurable tasks but struggle with the creativity required for open-ended scientific inquiry. Source: mittr