A new evaluation framework, reliability@k, addresses errors in the pass@k estimator used for coding agents. Correcting the metric by requiring independent rollouts instead of test-suite size reveals that current reported scores are inflated by up to 0.96 absolute points.
HOW THIS AFFECTS YOU
●
builderExpect more realistic performance benchmarks for the coding agents you integrate.
●
researcherYou must use reliability@k to avoid significantly overestimating agent performance.