Qiya Song
Papers
1
Total Citations
115
H-Index
1
About
Qiya Song is a leading researcher in multimodal machine learning and automatic speech recognition (ASR), with a focus on robust human-machine interaction in noisy environments. Their most influential work, the "Multimodal Sparse Transformer Network for Audio-Visual Speech Recognition" (2022, 115 citations), addresses a critical limitation of traditional ASR systems: severe performance degradation under external noise. By integrating visual lip movements with audio signals through a novel sparse attention mechanism, Song's architecture significantly improves recognition accuracy in challenging acoustic conditions—a breakthrough for applications in intelligent homes, autonomous driving, and servant robots. This contribution has been widely cited for its elegant fusion of transformer-based multimodal learning with computational efficiency. Song's research bridges the gap between theoretical deep learning and practical deployment, demonstrating how sparse attention can reduce computational overhead while enhancing robustness. Their work has inspired subsequent studies in audio-visual fusion and noise-robust ASR, establishing Song as a key innovator in making intelligent systems more reliable in real-world, acoustically cluttered environments.
Research Focus
Key Achievements
Top Papers
- 1Multimodal Sparse Transformer Network for Audio-Visual Speech Recognition115 citations · 2022