Warm-Starting Rubric-Based RL with On-Policy Distillation
October 1, 2026
This framework uses rubric-privileged on-policy distillation (RP-OPD) to provide dense token-level supervision for tasks lacking exact outcome verification. By matching a rubric-aware teacher's distributions before moving to RL, the model overcomes the sparse reward signals of open-ended tasks.
HOW THIS AFFECTS YOU
●
builderYou can use this technique to train models on complex, rubric-based tasks that are difficult to verify automatically.
●
researcherThis provides a two-stage method to stabilize RL for subjective or open-ended evaluation criteria.