RLVR Training Reduces Solution Space Coverage in Reasoning Models
August 28, 2026
Reinforcement learning with verifiable rewards (RLVR) improves pass@1 accuracy but contracts the policy's solution space, reducing solution coverage by up to 67% on the Countdown task. Analysis on Qwen2.5-3B shows this contraction limits the potential benefits of test-time scaling.
HOW THIS AFFECTS YOU
●
researcherYou must account for the trade-off between single-sample accuracy and the diversity of the reasoning trajectory when using RLVR.