Random-Reward RL Probes LLM Reachability Beyond Current Performance
October 2, 2026
Random-reward reinforcement learning identifies a model's reachability, or the potential performance ceiling achievable through further training. Testing OLMo checkpoints with identical 3.5% synthetic arithmetic accuracy revealed disparate potential, with one checkpoint reaching 8.5% and another 55% under the same correctness-rewarded RL.
HOW THIS AFFECTS YOU
●
researcherYou can use random rewards to probe the upper limits of a model's weight state during evaluation.