Reinforcement Learning with Verifiable Rewards (RLVR) can improve one-sample accuracy but decrease the model's ability to solve diverse problems at higher sampling rates. The researchers propose Per-Problem Base Anchoring (PBA) to prevent the loss of rare, correct trajectories during the RL process.
HOW THIS AFFECTS YOU
●
researcherThis provides a diagnostic framework for understanding why RLVR might shrink a model's reasoning boundary.