Agents fail open-ended research tasks in shadow evaluations
July 30, 2026
Agents tasked with replicating unpublished research papers over six days using thousands of dollars in compute failed to produce viable results. Authors of the original research unambiguously rejected the agent-generated papers, noting that while agents handle engineering tasks well, they struggle with hypothesis selection and navigating failing approaches.
HOW THIS AFFECTS YOU
●
researcherExpect current agentic workflows to remain limited to narrow, verifiable engineering tasks rather than open-ended scientific discovery.