OM-GRPO Decouples Reward Estimation from Policy Optimization in Label-Free RLVR
August 5, 2026
OM-GRPO prevents policy collapse in label-free Reinforcement Learning with Verifiable Rewards by masking gradients on answer spans. This method shifts optimization pressure from answer tokens to reasoning processes, using contrast-augmented rewards to refine estimation via pairwise comparisons without extra rollouts.
HOW THIS AFFECTS YOU
●
builderThis provides a pathway to scale RL training using model consensus rather than manual supervision.
●
researcherYou can use this to improve reasoning model training without relying on ground-truth labels.