Qiya Song

Hunan University

Papers

1

Total Citations

115

H-Index

1

About

Qiya Song is a leading researcher in multimodal machine learning and automatic speech recognition (ASR), with a focus on robust human-machine interaction in noisy environments. Their most influential work, the "Multimodal Sparse Transformer Network for Audio-Visual Speech Recognition" (2022, 115 citations), addresses a critical limitation of traditional ASR systems: severe performance degradation under external noise. By integrating visual lip movements with audio signals through a novel sparse attention mechanism, Song's architecture significantly improves recognition accuracy in challenging acoustic conditions—a breakthrough for applications in intelligent homes, autonomous driving, and servant robots. This contribution has been widely cited for its elegant fusion of transformer-based multimodal learning with computational efficiency. Song's research bridges the gap between theoretical deep learning and practical deployment, demonstrating how sparse attention can reduce computational overhead while enhancing robustness. Their work has inspired subsequent studies in audio-visual fusion and noise-robust ASR, establishing Song as a key innovator in making intelligent systems more reliable in real-world, acoustically cluttered environments.

Research Focus

Key Achievements

1
H-Index
1
Papers
115
Total Citations
115
Avg Citations/Paper
🏆 Most Cited Paper
Multimodal Sparse Transformer Network for Audio-Visual Speech Recognition
115 citations · 2022
📈 Most Prolific Year: 2022 (1 Papers)
🤝 Key Collaborators: 2
🏛 Institutions: Hunan University

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago