Statistical Limits of Extrapolated pass@k Evaluations
September 10, 2026
This analysis demonstrates that fixed-n rollout evaluations cannot reliably identify pass@k for values of k greater than the number of samples n. The research proves that even with large task counts, generic extrapolated tail constants and exponents remain unidentifiable without increasing the per-task rollout budget.
HOW THIS AFFECTS YOU
●
researcherYou should avoid trusting extrapolated pass@k metrics when your sample size n is smaller than the target k.