Counterfactual Sensitivity for Long-CoT Reasoning Credit Reallocation
July 29, 2026
This work addresses the limitation of uniform reward broadcasting in methods like GRPO by using counterfactual sensitivity to reallocate credit. By re-scoring trajectories under opposing outcome conditions, the method identifies which specific tokens actually drive the correctness of long-chain-of-thought reasoning.
HOW THIS AFFECTS YOU
●
researcherYou can improve RLVR training efficiency by moving beyond uniform token-level advantage broadcasting.