DataFlex-RL Benchmarks RLVR Data Policies for Qwen2.5-7B
September 4, 2026
DataFlex-RL evaluates reinforcement learning with verifiable rewards (RLVR) data policies using a GRPO recipe. Testing on Qwen2.5-7B-Base across 12 math, logic, and science benchmarks shows uniform GRPO improves domain-balanced accuracy by 7.76 percentage points, while adaptive mixtures failed to outperform fixed equal mixtures.
HOW THIS AFFECTS YOU
●
builderThis suggests uniform sampling might be a more robust baseline for RLVR than complex adaptive weighting schemes.
●
researcherYou can use this platform to standardize how you compare rollout-selection and reweighting methods in RLVR.