Balancing Reference Guidance and Free Generation in Reasoning RL
October 9, 2026
The study examines prefix continuation in Reinforcement Learning for reasoning models to find the optimal balance between reference guidance and free generation. It provides a closed-form KL divergence derivation to show how reference length affects the model's generated trajectory distribution.
HOW THIS AFFECTS YOU
●
researcherYou can optimize the length of reference prefixes in RL training to better match a model's natural distribution while ensuring correctness.