●builderThis offers a more scalable way to post-train multimodal models without relying solely on expensive online CoT sampling.
●researcherYou can improve RL sample efficiency for video models by integrating annotations directly into optimization targets.