Speech Emotion Recognition Using Hybrid VMD and EWT Based Cepstral Feature Extraction.
Abstract
Objective
Speech Emotion Recognition (SER) has gained significant research attention over the past three decades owing to its diverse real-world applications, including human-computer interaction, healthcare, call centers, automotive systems, education, and security. The primary goal of SER is to accurately identify human emotions from speech and enhance emotion classification performance. STUDY
Design
Recent studies have shown that signal decomposition-based feature extraction methods are more effective at capturing emotional cues than direct speech-level feature extraction, as different decomposition techniques can uncover complementary emotional information across the various frequency components of speech.
Methods
Motivated by this, the present study integrates two signal decomposition approaches, such as variational mode decomposition (VMD) and empirical wavelet transform (EWT), to decompose the speech signal frame (SSF) into multiple sub-signals, referred to as modes or intrinsic mode functions (IMFs). From these decomposed signals, features such as mel frequency cepstral coefficients (MFCC) and mel frequency magnitude coefficient (MFMC) are extracted. The VMD- and EWT-derived features are then combined to form the proposed variational mode empirical wavelet transform-based mel frequency cepstral and magnitude coefficient (PVEWMFCMC) features, which are utilized for emotion classification using a Deep Neural Network (DNN) classifier.
Results
Experimental evaluations demonstrate that the proposed PVEWMFCMC features, combined with a DNN model, achieved speaker-dependent classification accuracies of 91.76%, 86.92%, and 81.53% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively. To further evaluate the generalization capability of the proposed approach, speaker-independent evaluation using the Leave-One-Speaker-Out (LOSO) protocol was also performed, achieving accuracies of 74.59%, 58.43%, and 56.79% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively.