STATERA Predicts Hidden Mass via Temporal Tubelet Mixing
October 2, 2026
STATERA adapts a pretrained V-JEPA backbone with a temporal tubelet mixer to estimate center-of-mass (CoM) from monocular video. In simulation, the model reduced normalized CoM error from 41.7% to 25.2% compared to DINOv2.
HOW THIS AFFECTS YOU
●
builderThis provides a pathway for extracting physical properties like mass from standard video inputs.
●
researcherThe work highlights how temporal mixers can bridge the gap between appearance and physical property inference.