Home /Research /Deep Neural Network for visual Emotion Recognition based on ResNet50 using Song-Speech characteristics
LEARNING

Deep Neural Network for visual Emotion Recognition based on ResNet50 using Song-Speech characteristics

Souha Ayadi, Zied Lachiri

Year
2022
Citations
11

Abstract

Visual emotion recognition is a very large field. It plays a very important role in different domains such as security, robotics, and medical tasks. The visual tasks could be either image or video. Unlike the image processing, the difficulty of video processing is always a challenge due to changes in information over time variation. Significant performance improvements when applying deep learning algorithms to video processing. This paper presents a deep neural network based on ResNet50 model. The latter is conducted on the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) due to the variance of the nature of the data exists which is speech and song. The choice of ResNet model is based on the ability of facing different problems such as of vanishing gradients, the performing stability offered by this model, the ability of CNN for feature extraction which is considered to be the base architecture for ResNet, and the ability of improving the accuracy results and minimizing the loss. The achieved results are 57.73% for song and 55.52% for speech. Results shows that the Resnet50 model is suitable for both speech and song while maintaining performance stability.

Keywords

Computer scienceSpeech recognitionArtificial intelligenceFeature extractionArtificial neural networkField (mathematics)Deep learningSpeech processingStability (learning theory)Feature (linguistics)

Related papers

Browse all LEARNING papers