World in World: Training-Free Video World Model Control
September 9, 2026
World in World provides a training-free inference-time interface for autoregressive video world models. It converts camera and time-labeled evidence into visual states via the model's native self-attention to enable flexible exploration.
HOW THIS AFFECTS YOU
●
researcherYou can achieve interactive video exploration without requiring task-specific module training.
●
designerThis enables more precise control over viewpoints and time in generated video sequences.