RL Updates in VLA Models are Low-Rank and Concentrated
September 27, 2026
Reinforcement learning in vision-language-action models like π_{0.5} induces low-rank parameter updates concentrated in the action expert's Timestep Modules. These small modules capture a disproportionate share of performance gains during RL post-training.
HOW THIS AFFECTS YOU
●
builderYou can potentially reduce RL fine-tuning costs by parameter-efficiently targeting the action expert's timestep components.
●
researcherRL-based policy refinement can be targeted more efficiently by focusing on specific timestep modules.