New Scaling Law Connects Reward Model Data to Optimization Budgets
October 1, 2026
Performance in reward optimization scales as $\Theta(\sqrt{\min\{\log(M),K\}})$, where $M$ is the number of training comparisons and $K$ is the KL-divergence budget. This law quantifies how reward hacking occurs when optimization effort exceeds the informational capacity of the proxy reward model.
HOW THIS AFFECTS YOU
●
builderYou can better estimate the data requirements needed to prevent reward hacking during policy optimization.
●
researcherYou can use this theoretical grounding to predict the point of diminishing returns in RLHF pipelines.