Both agent-written papers were rejected
Princeton researchers gave agents six days, $3,000 in credits and unpublished research questions. The original authors turned down the results.
2 minMain AI Hub
A team at Princeton led by Peter Kirgis and Sayash Kapoor, working across several institutions, set AI agents to answer unpublished research questions taken from papers submitted to NeurIPS 2026. The agents ran on Claude Opus 4.8 through the open-source OpenClaw software, with six days, $3,000 in API credits, GPU access, virtual machines and open web access.
Two papers came out of it. Both were rejected by the human authors of the original work.
Where they failed is specific
The agents handled the engineering competently. What they lacked was judgement: they did not explore a wide enough range of ideas, committed to approaches that were failing rather than abandoning them, and could not fundamentally rethink a strategy once it was underway. They also struggled to act on feedback and to manage their own budget of compute time and tokens.
The finding bears directly on the recursive self-improvement timelines that circulate in the industry. A system that cannot recognise a dead end and change direction is not going to improve itself by producing better research, whatever its scores on bounded tasks.
The shape of the result is familiar and worth naming: strong on narrow, measurable work, slow on open-ended work where the task includes deciding what the task is. Two rejected papers is a small sample, and the study measures one model on one harness — but the failure modes it lists are mechanisms, not scores, and mechanisms generalise further than benchmarks do.
Retold from MIT Technology Review. This is a summary in our own words; follow the link for the original reporting.