FLAT Framework Synchronizes Multimodal Representation and Generative Decoding
September 14, 2026
FLAT optimizes a shared multimodal encoder alongside text-to-image and image-to-text decoders through joint contrastive and bidirectional generative objectives. This produces 1D flexible-length aligned tokens that allow direct, linear interpolation between modalities for improved retrieval and generation performance.
HOW THIS AFFECTS YOU
●
builderYou can use directly consumable transmodal tokens to streamline multimodal retrieval and generation pipelines.
●
researcherYou can bypass the bottleneck of frozen embeddings by training joint representation-generative models.