Skip to content

Speech Emotion Recognition Using Hybrid VMD and EWT Based Cepstral Feature Extraction.

Sep 2026 · Journal of Voice · 0 citations · 47 references
Medicine

Abstract

Objective

Speech Emotion Recognition (SER) has gained significant research attention over the past three decades owing to its diverse real-world applications, including human-computer interaction, healthcare, call centers, automotive systems, education, and security. The primary goal of SER is to accurately identify human emotions from speech and enhance emotion classification performance. STUDY

Design

Recent studies have shown that signal decomposition-based feature extraction methods are more effective at capturing emotional cues than direct speech-level feature extraction, as different decomposition techniques can uncover complementary emotional information across the various frequency components of speech.

Methods

Motivated by this, the present study integrates two signal decomposition approaches, such as variational mode decomposition (VMD) and empirical wavelet transform (EWT), to decompose the speech signal frame (SSF) into multiple sub-signals, referred to as modes or intrinsic mode functions (IMFs). From these decomposed signals, features such as mel frequency cepstral coefficients (MFCC) and mel frequency magnitude coefficient (MFMC) are extracted. The VMD- and EWT-derived features are then combined to form the proposed variational mode empirical wavelet transform-based mel frequency cepstral and magnitude coefficient (PVEWMFCMC) features, which are utilized for emotion classification using a Deep Neural Network (DNN) classifier.

Results

Experimental evaluations demonstrate that the proposed PVEWMFCMC features, combined with a DNN model, achieved speaker-dependent classification accuracies of 91.76%, 86.92%, and 81.53% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively. To further evaluate the generalization capability of the proposed approach, speaker-independent evaluation using the Leave-One-Speaker-Out (LOSO) protocol was also performed, achieving accuracies of 74.59%, 58.43%, and 56.79% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.