ReMoMask-2: Structure-Aware RAG for Text-to-Motion Generation
September 7, 2026
ReMoMask-2 addresses the representation gap in retrieval-augmented motion generation by coupling Hierarchical Bidirectional Momentum contrastive learning with Semantic Spatial-Temporal Attention. This aligns global and part-level features with text to capture human motion topology.
HOW THIS AFFECTS YOU
●
builderYou can generate more anatomically and spatially coherent human motions from text via RAG.
●
researcherThe hierarchical approach offers a method to bridge the gap between semantic text and latent motion spaces.