DeltaML-Bench Evaluates Agents on Real-World Research Repositories
August 21, 2026
DeltaML-Bench evaluates ML agents on 48 tasks requiring them to improve published baselines within imperfect open-source repositories. Using search-based ARG scaffolding, GPT-5's success rate on these tasks rose from 9.4% to 49.0% under long-horizon allocations.
HOW THIS AFFECTS YOU
●
builderUse this benchmark to test if your autonomous agents can actually fix and improve messy research codebases.
●
researcherThis provides a more realistic evaluation framework for agents operating under compute and repository constraints.