Hierarchical World Models for Symbolic Music Generation
August 4, 2026
A 2.55M-parameter Swin V2 encoder uses JEPA-style objectives to learn hierarchical representations of MIDI piano-roll images without labels. The model captures musical properties, from fine-grained note density to coarse-grained phrase boundaries, across different temporal scales.
HOW THIS AFFECTS YOU
●
researcherThis demonstrates the effectiveness of self-supervised world models for symbolic music.
●
designerThis may lead to more coherent and musically intelligent co-creation tools.