RL^2-VLA Adaptive Inference-Time Steering for VLA Models
July 29, 2026
RL^2-VLA is an adaptive inference-time steering framework that applies reinforcement learning to Vision-Language-Action (VLA) latents. Unlike static intervention methods, it scales test-time compute by conditioning steering on whether the base policy is likely to succeed or fail.
HOW THIS AFFECTS YOU
●
builderYou can improve VLA performance on out-of-domain tasks using lightweight offline RL on latents.
●
researcherThis provides a way to mitigate correlated failure modes in test-time scaling for robotics.