GRENZE International Journal of Engineering and Technology
Vol. 10
(2024), Issue 1
End to End Speech Emotion Recognition
Authors
Srivatsa V, Vivek B V, Sumana M, Vashista A N
Abstract
The proposed model combines machine learning, artificial neural networks (ANN), and natural language processing (NLP) techniques for a robust analysis of emotions in speech. The process begins by extracting relevant features from the audio signal, such as Mel Frequency Cepstral Coefficients (MFCC), which capture the spectral characteristics of the speech signal. These features are used to train a feedforward artificial neural network (ANN) to classify the emotions in the speech, with the model optimized using techniques like dropout and adaptive learning rates to prevent overfitting and improve generalization performance. In addition to the audio features, the model incorporates natural language processing (NLP) techniques to analyse the content of the speech. It uses speech-to-text conversion to obtain the transcript of the speech, and then applies sentiment interpretation using TextBlob to calculate the sentiment polarity of the text. This information is combined with the emotion predicted by the ANN model to create a more effective emotion categorization, taking into account both the acoustic and linguistic aspects of the speech
Pages:
589 - 597