SISER Improves Speech Emotion Recognition via Adversarial Training
September 4, 2026
SISER combines wav2vec 2.0 feature encoding with ECAPA-TDNN speaker discrimination using entropy-based adversarial training. On the IEMOCAP dataset, the method achieved a Unweighted Accuracy of 60.63%, outperforming the wav2vec 2.0 baseline of 56.46%.
HOW THIS AFFECTS YOU
●
researcherThe integration of ECAPA-TDNN as a discriminator provides a stronger signal for suppressing speaker identity in emotion recognition tasks.