CorrGRPO Normalizes Rewards for Multi-Objective Learning
September 28, 2026
CorrGRPO improves Group Relative Policy Optimization for multi-reward scenarios by normalizing pairwise reward covariances into Pearson correlations. This prevents large-scale correlated rewards from dominating the advantage signal and suppressing smaller-scale reward components.
HOW THIS AFFECTS YOU
●
researcherYou can use this to more effectively train reasoning models with complex, multi-objective reward functions.