Benchmarking EEG Foundation Models for Clinical Robustness
July 28, 2026
Evaluation of six EEG foundation models, including LaBraM and EEGMamba, reveals significant performance gaps in clinical tasks. Frozen REVE embeddings achieved only 0.568 AUROC on Korean dementia tasks compared to 0.769 for classical features, while dataset identity was nearly perfectly decodable (AUROC ~1.0).
HOW THIS AFFECTS YOU
●
researcherYou should be skeptical of foundation model performance on clinical datasets due to high signal leakage from dataset identity.
●
healthCurrent EEG foundation models may lack the necessary robustness for reliable clinical diagnostic tools.