DAPO System Achieves 50 Points on AIME 2024 via RL
September 21, 2026
DAPO is an open-source system designed for large-scale LLM reinforcement learning. Using the Qwen2.5-32B base model, the system achieved a score of 50 on the AIME 2024 benchmark.
HOW THIS AFFECTS YOU
●
builderYou can leverage these RL techniques to improve the reasoning capabilities of smaller, more efficient models.
●
researcherYou can use this open-source framework to implement large-scale RL for mathematical reasoning tasks.