Home /Research /Multi-Modal Emotion Recognition Network with Balanced Audio-Visual Feature Extraction
OTHER

Multi-Modal Emotion Recognition Network with Balanced Audio-Visual Feature Extraction

Zeyu Chen, Yiming Wu, Ronghui Cao

Year
2024
Citations
1

Abstract

Video and audio are ways for humans to perceive the world beyond language. It is meant to enable robots to imitate and recognize human emotional expressions. However, most current audio-visual emotion analysis models tend to extract deep features from only one modality. In contrast, the other modality plays a supporting role and cannot fully extract deep features from both modalities. This article proposes a balanced audio-visual emotion analysis model. Specifically, EfficientNet and Wav2vec 2.0 are used for visual and auditory modality feature extraction, respectively, ensuring that indepth features can be extracted from both modalities. Secondly, we use Transformer as the decision-level fusion operator, exchanging information between the two modalities. We verified our model on the RAVDESS dataset and achieved a Top1 accuracy of 88.54%, surpassing audio-visual emotion analysis models of the auxiliary type.

Keywords

Feature extractionComputer scienceModalEmotion recognitionAudio visualSpeech recognitionArtificial intelligenceFeature (linguistics)Pattern recognition (psychology)Multimedia

Related papers

Browse all OTHER papers