AGRPO algorithm improves reasoning in diffusion language models
July 24, 2026
Amortized Group Relative Policy Optimization (AGRPO) optimizes individual denoising steps in diffusion large language models rather than full sequences. This method achieved accuracy gains of +59.4% on Countdown and +69.7% on Sudoku compared to base models.
HOW THIS AFFECTS YOU
●
researcherYou can apply this timestep estimation scheme to improve alignment in non-autoregressive model training.