Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
August 14, 2026
Evaluating seven frontier models across 36 long-horizon tasks reveals that current agents function as engineering optimizers rather than autonomous researchers. A new framework using rule-based metrics for solution framing and feedback control shows that agents struggle with performance consistency and effective experience reuse during extended experimentation.