CoRT: Token-Level Credit Weighting for Rubric-Guided Policy Optimization
July 29, 2026
CoRT enables token-level reward allocation in GRPO-style pipelines using counterfactual replay. By contrasting log-likelihoods between rubric-conditioned and criteria-free prompts, it maps rubric dependence to specific token spans rather than using uniform response-level scalars.
HOW THIS AFFECTS YOU
●
builderThis allows you to train models that strictly follow complex formatting and semantic rubrics.
●
researcherYou can move beyond scalar rewards to more granular, token-level reinforcement learning.