Protocol Defects Inflate AutoML Performance Metrics in Short-Budget Tests
August 10, 2026
Analysis of Orcetra shows it appeared to outperform FLAML and AutoGluon on 513 datasets due to test-set leakage and unenforced time budgets. The study demonstrates that scoring candidates on the test split and ignoring budget limits during execution creates deceptive performance leads.
HOW THIS AFFECTS YOU
●
researcherYou should verify that AutoML benchmarks use strict validation splits and enforced wall-clock limits to avoid inflated results.