Facial Expression Recognition Based on Multi-Scale Convolutional Vision Transformer
Cheng-Shan Jiang, Zhentao Liu
- 发表年份
- 2022
- 引用次数
- 6
摘要
Facial expression recognition (FER) could endow artificial intelligence devices such as service robots with better understanding of emotional state of human beings, which facilitates the experience of human-computer interaction more harmonious and natural. The convolution operation will be limited by the receptive field, and the extracted features are local, so it is hard to understand and learn the facial expression information from a global point of view. In this paper, a Multi-Scale Convolutional Vision Transformer (MSC-ViT) is proposed for FER. It replaces the linear embedding with Convolutional Tokenization, and uses the Multi-Scale Convolutional Position Mapping (MSCPM) to obtain the multi-scale feature information of each facial expression image patches, and carries on the information integration and feature learning from the global perspective of view. We verify the performance of the MSC-ViT on the RaFD data set, and the recognition accuracy is 98.26%.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991