Speech Emotion Recognition using Convolutional Neural Networks and Recurrent Neural Networks with Attention Model
Abstract
Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outlines a study in which a data set of speech signals—containing semantically neutral recordings of professional actors portraying eight different emotions by altering their voices—is tested. Low-level descriptors (LLDs) were extracted as static subfeatures from the speech signals and classified with support vector machines (SVMs) and convolutional neural networks (CNNs). Furthermore, mel-spectrograms were extracted and fed into a CNN as 3D image vectors for classificaiton. The resulting accuracies from both the SVMs and CNNs for LLDs proved to be better than human raters from a previous study, with the CNN classifyer achieving the highest accuracy. The CNN for mel-spectrograms failed to achieve similar results, which is explained by lack of computational resources in the experiment setup.