Marketing narratives from Anthropic and OpenAI pitching autonomous AI research as an imminent reality have run into hard empirical ground. A joint study from Princeton and the UK AI Security Institute evaluated leading reasoning models against unpublished NeurIPS submissions using a benchmark design termed Shadow Evaluation. Rather than benchmarking synthetic sub-tasks or relying on automated review, researchers handed the agents core research questions and had the original paper authors score the resulting drafts. The trials granted frontier models six days of compute, a $3,000 API budget, GPU access, and full web toolchains.
The human authors unanimously rejected the AI-generated papers, handing down a "Strong Reject" over unreadable prose, poorly motivated experiments, and zero genuine conceptual contributions. While the models executed routine engineering baselines without human intervention, they collapsed when tasked with scientific judgment.
For technical leaders and R&D managers, the takeaway is clear: current reasoning agents operate as capable research assistants for debugging, data wrangling, and mechanical workflows, but they remain incapable of formulating novel scientific hypotheses. Managing code execution is a far cry from evaluating scientific merit, keeping the creative direction of research firmly in human hands.