Deep Neural Network for visual Emotion Recognition based on ResNet50 using Song-Speech characteristics
Souha Ayadi, Zied Lachiri
- Year
- 2022
- Citations
- 11
Abstract
Visual emotion recognition is a very large field. It plays a very important role in different domains such as security, robotics, and medical tasks. The visual tasks could be either image or video. Unlike the image processing, the difficulty of video processing is always a challenge due to changes in information over time variation. Significant performance improvements when applying deep learning algorithms to video processing. This paper presents a deep neural network based on ResNet50 model. The latter is conducted on the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) due to the variance of the nature of the data exists which is speech and song. The choice of ResNet model is based on the ability of facing different problems such as of vanishing gradients, the performing stability offered by this model, the ability of CNN for feature extraction which is considered to be the base architecture for ResNet, and the ability of improving the accuracy results and minimizing the loss. The achieved results are 57.73% for song and 55.52% for speech. Results shows that the Resnet50 model is suitable for both speech and song while maintaining performance stability.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002