NeoMME: Multimodal-Native Multilingual Foundation Encoder
August 30, 2026
NeoMME provides 260M and 800M-parameter bidirectional encoders that process multilingual text and raw image patches in a single Transformer, supporting a 16,384-token context.
HOW THIS AFFECTS YOU
●
builderThese models offer a more efficient alternative to repurposing large generative VLMs for non-generative retrieval tasks.
●
researcherThe architecture utilizes a masked discrete-diffusion text objective conditioned on image patches.