MOSS-VL Enables Real-Time Interaction via Gated Cross-Attention
August 14, 2026
MOSS-VL is an open vision-language model family designed for real-time interaction, using gated cross-attention to allow the language decoder to perceive incoming video frames during generation. It outperforms existing open-source streaming models on three of four streaming benchmarks and leads in temporal-reasoning video sets.
HOW THIS AFFECTS YOU
●
builderYou can use this architecture to build low-latency agents that perceive and speak simultaneously.
●
researcherThe staged curriculum for real-time training offers a new framework for multimodal temporal reasoning.