首页 /研究 /Multi-Modal Emotion Recognition Network with Balanced Audio-Visual Feature Extraction
OTHER

Multi-Modal Emotion Recognition Network with Balanced Audio-Visual Feature Extraction

Zeyu Chen, Yiming Wu, Ronghui Cao

发表年份
2024
引用次数
1

摘要

Video and audio are ways for humans to perceive the world beyond language. It is meant to enable robots to imitate and recognize human emotional expressions. However, most current audio-visual emotion analysis models tend to extract deep features from only one modality. In contrast, the other modality plays a supporting role and cannot fully extract deep features from both modalities. This article proposes a balanced audio-visual emotion analysis model. Specifically, EfficientNet and Wav2vec 2.0 are used for visual and auditory modality feature extraction, respectively, ensuring that indepth features can be extracted from both modalities. Secondly, we use Transformer as the decision-level fusion operator, exchanging information between the two modalities. We verified our model on the RAVDESS dataset and achieved a Top1 accuracy of 88.54%, surpassing audio-visual emotion analysis models of the auxiliary type.

关键词

Feature extractionComputer scienceModalEmotion recognitionAudio visualSpeech recognitionArtificial intelligenceFeature (linguistics)Pattern recognition (psychology)Multimedia

相关论文

查看 OTHER 分类全部论文