Papers
1
Total Citations
414
H-Index
1
About
Yueming Ding is a leading researcher in multimodal artificial intelligence, with a primary focus on visual question answering (VQA), image-text fusion, and multiscale feature extraction. Ding’s most influential contribution is the development of novel architectures that seamlessly integrate visual and textual information, enabling machines to reason about images in response to natural language queries. Their landmark 2023 paper, “Multiscale Feature Extraction and Fusion of Image and Text in VQA,” has garnered over 414 citations, underscoring its impact on advancing intelligent systems for visual assistance, automated security surveillance, and human-robot interaction. This work introduced innovative techniques for capturing and combining features across multiple scales, significantly improving the accuracy and robustness of VQA models. Ding’s research bridges the gap between computer vision and natural language processing, pushing the boundaries of how AI understands and interacts with the visual world. Their contributions are foundational for developing more intuitive and context-aware AI assistants, making Ding a key figure in the evolution of multimodal learning and intelligent interaction systems.
Research Focus
Key Achievements
Top Papers
- 1Multiscale Feature Extraction and Fusion of Image and Text in VQA414 citations · 2023