Aug 2026· PLoS ONE· Vol 21· 0 citations· 54 references
Medicine
TL;DR
The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit, and improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.
Abstract
Speech Emotion Recognition (SER) is an important component of human-centered intelligent systems, yet robust performance remains challenging when speaker identities differ between training and testing. This study presents a protocol-aware and reproducible comparison on the CREMA-D corpus using three pipelines: (i) a classical MFCC-based Support Vector Machine (SVM), (ii) a log-mel Convolutional Neural Network (CNN), and (iii) a lightweight hybrid model that concatenates handcrafted acoustic descriptors with CNN-derived embeddings and uses an SVM classifier. The methodological contribution is not a new standalone classifier; it is the controlled integration of identical preprocessing, random and actor-wise evaluation, five-seed robustness reporting, class-wise error analysis, and CPU-oriented deployment within one experimental framework. The Hybrid approach achieves the best overall performance, obtaining 62.03% ± 0.94% Macro-F1 on the random split and 58.09% ± 1.36% on the actor-wise split, outperforming MFCC+SVM (55.72% ± 0.95% and 52.05% ± 1.95%) and the log-mel CNN (50.20% ± 1.51% and 42.68% ± 1.84%). A Streamlit interface supports WAV upload, live prediction, and export of per-seed confusion matrices and summary figures. The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.
Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...
G. Tomas, S. Weinzierl, Athanasios Lykartsis· Proceedings of 2019 the 9th...· 2 citations
A bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture and the Multimodal EmotionLines Dataset is proposed, suggesting that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER.
This work proposes SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme.
Human speech carries not only words but also rich paralinguistic signals that reflect emotional states. Despite significant progress in automatic speech emotion recognition (SER), many existing models still struggle to fully exploit vocal information. In this paper, we propose a novel deep hybrid model with three paral...
Sara Rekkal, Kahina Rekkal, Adem Abdelhamid Boulhadid· International journal of ele...· 0 citations
This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.
Muhammad Sheraz, Adil Majeed, Shehzad Khalid et al.· Computer Modeling in Enginee...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.