LeakScale Quantifies Causal Impact of Benchmark Contamination
September 24, 2026
LeakScale is an interventional framework that estimates the exact performance gain attributable to benchmark exposure by using fresh, executable tasks with private information. Testing across 2,048 families shows exposure improves accuracy by 7.17 to 27.31 percentage points, separating mere provenance from true causal effect.
HOW THIS AFFECTS YOU
●
researcherYou can now quantify exactly how much training data leakage inflates your model's reported scores.
●
policyThis provides a more rigorous metric for evaluating whether models are truly capable or just memorizing evaluation sets.