Introduces ZCPO to address failure modes in group policy optimization caused by KL regularization. It uses conditional KL to calibrate within-group reward coefficients, preventing issues like KL growth with response length.
HOW THIS AFFECTS YOU
●
researcherYou can implement more stable RL training when optimizing group policies.