Skip to content
Open access

Hybrid CNN-embedding fusion with MFCC-SVM for speech emotion recognition: Random vs actor-wise evaluation on CREMA-D

Aug 2026 · PLoS ONE · Vol 21 · 0 citations · 54 references
Medicine

TL;DR

The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit, and improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.

Abstract

Speech Emotion Recognition (SER) is an important component of human-centered intelligent systems, yet robust performance remains challenging when speaker identities differ between training and testing. This study presents a protocol-aware and reproducible comparison on the CREMA-D corpus using three pipelines: (i) a classical MFCC-based Support Vector Machine (SVM), (ii) a log-mel Convolutional Neural Network (CNN), and (iii) a lightweight hybrid model that concatenates handcrafted acoustic descriptors with CNN-derived embeddings and uses an SVM classifier. The methodological contribution is not a new standalone classifier; it is the controlled integration of identical preprocessing, random and actor-wise evaluation, five-seed robustness reporting, class-wise error analysis, and CPU-oriented deployment within one experimental framework. The Hybrid approach achieves the best overall performance, obtaining 62.03% ± 0.94% Macro-F1 on the random split and 58.09% ± 1.36% on the actor-wise split, outperforming MFCC+SVM (55.72% ± 0.95% and 52.05% ± 1.95%) and the log-mel CNN (50.20% ± 1.51% and 42.68% ± 1.84%). A Streamlit interface supports WAV upload, live prediction, and export of per-seed confusion matrices and summary figures. The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.

Read PDF

Similar papers

Open access 2019

Speech Emotion Recognition using Convolutional Neural Networks and Recurrent Neural Networks with Attention Model

Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...

G. Tomas, S. Weinzierl, Athanasios Lykartsis · 2 citations
Open access Aug 2026

Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion

A bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture and the Multimodal EmotionLines Dataset is proposed, suggesting that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER.

Tze-Syn Yap, Lee-Yeng Ong · 0 citations
Open access Sep 2026

A novel deep hybrid parallel model of CNN- BiLSTM-Transformer for speech emotion recognition

Human speech carries not only words but also rich paralinguistic signals that reflect emotional states. Despite significant progress in automatic speech emotion recognition (SER), many existing models still struggle to fully exploit vocal information. In this paper, we propose a novel deep hybrid model with three paral...

Sara Rekkal, Kahina Rekkal, Adem Abdelhamid Boulhadid · 0 citations
Open access 2026

Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations

This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.

Muhammad Sheraz, Adil Majeed, Shehzad Khalid et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.