RLVR Causes Verifier-Induced Support Reshaping in Optimization
July 30, 2026
On-policy reinforcement learning with verifiable rewards (RLVR) can narrow the distribution of successful trajectories, making certain behaviors too rare to reinforce. In experiments with Qwen3-8B-Base, Math-RLVR improved instruction-following but reduced the diversity of successful responses under repeated sampling.
HOW THIS AFFECTS YOU
●
researcherYou should be aware that verifiable rewards can unintentionally collapse the support of your training distribution.