AutoResearchEval evaluates agentic scaffolds across 100 tasks spanning the full research lifecycle, from ideation to review. The study analyzes 800 trajectories to categorize failure modes in autonomous scientific discovery processes.
HOW THIS AFFECTS YOU
●
builderThis reveals why current agentic scaffolds fail at complex, multi-stage reasoning required for real-world research.
●
researcherYou can use this dataset to identify specific breakdown points in the end-to-end scientific discovery loop.