Applying the Muon optimizer to hidden weight matrices in sparse-reward agentic RL (using GiGPO on Qwen2.5-0.5B) increased validation success from 0.290 to 0.546. This outperforms AdamW controls in RL post-training settings.
HOW THIS AFFECTS YOU
●
researcherYou can improve reinforcement learning agent performance by swapping AdamW for Muon on hidden weights.