CoTinyVLA Distills Chain-of-Thought into a 0.9B Parameter VLA Model
July 29, 2026
CoTinyVLA uses a Qwen3.5-0.8B backbone to achieve 90.8% spatial robustness on the LIBERO-Plus benchmark. It utilizes hierarchical chain-of-thought distillation from a 35B teacher and dual-view temporal inputs to enable high-performance robotics on embedded hardware.
HOW THIS AFFECTS YOU
●
builderYou can deploy high-performance vision-language-action models on resource-constrained robotic hardware.
●
researcherThis demonstrates the effectiveness of hierarchical CoT distillation for scaling down VLA models.