Ambitions to automate AI research hit a wall in a new test led by Princeton’s Sayash Kapoor. An agentic system built on Claude Opus 4.8, equipped with tools for literature review, experiment execution and self-critique, was given six days and $3,000 in compute per task to extend ideas from two NeurIPS submissions. While the system ran hundreds of experiments and caught many of its own errors, original authors graded the outputs just 2/6 and 1/6, citing premature commitment to weak hypotheses and insufficient backtracking. The findings suggest AI can speed routine research chores but isn’t close to autonomously delivering significant advances, underscoring limits of peer review as a quality bar and the need for tighter evaluation of “AI scientists.”
Related articles:
ReAct: Synergizing Reasoning and Acting in Language Models
Toolformer: Language Models Can Teach Themselves to Use Tools
Reflexion: Language Agents with Verbal Reinforcement Learning
Self-Refine: Iterative Refinement with Self-Feedback
Voyager: An Open-Ended Embodied Agent in Minecraft



























