Skip to content

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

This work proposes SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme.

Abstract

Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches address speaker variability, they fall short in leveraging powerful pre-trained representations. We propose SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme. wav2vec 2.0 provides rich self-supervised representations that alleviate dependency on large labeled datasets, while ECAPA-TDNN enables suppression of speaker identity via a stronger adversarial signal than shallow classifiers. Evaluated on IEMOCAP, SISER achieves a UA of 60.63%, outperforming the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%), with ablation emphasizing that the choice of speaker classifier architecture is a key factor.

View source

Similar papers

Review Sep 2026

Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially,...

Xin-Yuan Qian, Yang Zhou, Zi-Yang Jiang et al. · 0 citations
Open access Aug 2026

Hybrid CNN-embedding fusion with MFCC-SVM for speech emotion recognition: Random vs actor-wise evaluation on CREMA-D

The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit, and improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit...

Parveen Kumari, Yogita Yashveer Raghav, Vimmi Kochher et al. · 0 citations
Open access Sep 2026

Emotion-Robust Speaker Recognition Through Synthetic Emotional Spectrogram Augmentation

Speaker recognition models are typically enrolled using neutral speech, yet real users rarely speak under emotionally neutral conditions. Emotion alters prosody, spectral structure, articulation, and speaking rate, shifting utterances away from the acoustic distribution observed during enrolment. Conventional augmentat...

E. Jakubcová, Maroš Jakubec · 0 citations
#natural language process... Preprint Sep 2026

Emotion as a Distribution: Joint Valence-Arousal Probability Learning for Speaker-Independent Multimodal Emotion Recognition

This text+speech system, alongside its categorical decision, emits a $9\times9$ probability matrix over the Valence-Arousal plane, trained with a two-dimensional Gaussian soft target under a Kullback-Leibler/cross-entropy objective, aimed at counseling support.

Ting Lin, Wen-Ren Yang, Kuan-Wei Chen · 0 citations
#artificial intelligence Preprint Sep 2026

From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models

Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted...

He-Zhao Zhang, Thomas Hain · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.