Addressing Advantage Scale Imbalance in Group-Relative Optimization
September 18, 2026
This paper diagnoses why standard GRPO can amplify tiny reward gaps, leading to KL calibration issues in RLVR. The proposed MaxNorm-AC interface provides a three-way calibration to filter sub-resolution jitter and bound the influence of small cardinal gaps.
HOW THIS AFFECTS YOU
●
researcherThis provides a mathematical method to ensure more stable and well-calibrated RL training when using group-relative rewards.