Taxonomy of Rollout Efficiency in Reasoning-Oriented RL
September 23, 2026
Reasoning-oriented reinforcement learning shifts significant training costs to the rollout phase during trajectory generation. This survey classifies existing efficiency techniques by mechanism and bottleneck to address data freshness, consistency, and statistical validity.
HOW THIS AFFECTS YOU
●
builderThis highlights the primary cost bottleneck you will face when training reasoning models.
●
researcherYou can use this taxonomy to identify gaps in existing rollout optimization literature.