Source Observation During Pretraining Improves Reward Adaptation in LLMs
October 1, 2026
This study demonstrates that providing task-independent source observations during pretraining allows models to better utilize rewards for adaptation. Experiments with Qwen2.5 show that models with source supervision achieve 82.61% success in sequential tasks compared to 44.15% without it.
HOW THIS AFFECTS YOU
●
researcherThe findings suggest that pretraining data composition is critical for how effectively a model learns from later RLHF.