CoVer Framework Uses Information-Gain Rewards for Code Generation
September 21, 2026
CoVer implements a single-policy GRPO framework to solve permissiveness collapse and concentration bias in self-play code generation. It utilizes information-gain rewards to score self-generated tests based on their mutual information with ground-truth correctness signals.
HOW THIS AFFECTS YOU
●
researcherThis addresses fundamental pathologies in training coder-verifier models via reinforcement learning.