Multimodal Mamba-ViT Foundation Model for EEG Representation Learning
July 24, 2026
This EEG foundation model integrates a Mamba-based raw signal encoder, a Vision Transformer for time-frequency data, and a text encoder into a shared embedding space. It employs masked modeling and cross-view contrastive alignment to learn seizure-relevant representations without labeled data.
HOW THIS AFFECTS YOU
●
researcherThe architecture combines Mamba and ViT to handle high-dimensional time-series and frequency data in a single backbone.
●
healthThis approach could lead to more generalizable diagnostic tools for epilepsy across diverse patient datasets.