Papers

1

Total Citations

414

H-Index

1

About

Yueming Ding is a leading researcher in multimodal artificial intelligence, with a primary focus on visual question answering (VQA), image-text fusion, and multiscale feature extraction. Ding’s most influential contribution is the development of novel architectures that seamlessly integrate visual and textual information, enabling machines to reason about images in response to natural language queries. Their landmark 2023 paper, “Multiscale Feature Extraction and Fusion of Image and Text in VQA,” has garnered over 414 citations, underscoring its impact on advancing intelligent systems for visual assistance, automated security surveillance, and human-robot interaction. This work introduced innovative techniques for capturing and combining features across multiple scales, significantly improving the accuracy and robustness of VQA models. Ding’s research bridges the gap between computer vision and natural language processing, pushing the boundaries of how AI understands and interacts with the visual world. Their contributions are foundational for developing more intuitive and context-aware AI assistants, making Ding a key figure in the evolution of multimodal learning and intelligent interaction systems.

Research Focus

Key Achievements

1
H-Index
1
Papers
414
Total Citations
414
Avg Citations/Paper
🏆 Most Cited Paper
Multiscale Feature Extraction and Fusion of Image and Text in VQA
414 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 5
🏛 Institutions: University of Electronic Science and Technology of China

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 11 days ago