Xinxin Xu
Papers
1
Total Citations
5
H-Index
1
About
Xinxin Xu is a leading researcher at the forefront of multimodal 3D scene understanding, a field that bridges computer vision, natural language processing, and robotics. Their seminal survey, "A survey of language-grounded multimodal 3D scene understanding" (2025), has already garnered 5 citations, establishing a foundational roadmap for integrating linguistic cues with 3D spatial data. Xu’s major contributions lie in developing frameworks that enable machines to interpret and interact with complex 3D environments through natural language, advancing embodied AI and autonomous systems. By synthesizing disparate approaches—from point cloud processing to vision-language models—Xu has clarified key challenges and opportunities in grounding language in physical space, influencing how researchers design systems for tasks like scene description, navigation, and object manipulation. This work has quickly become a go-to reference for students and scholars exploring the intersection of language and 3D perception. With a growing citation impact and a knack for identifying pivotal research directions, Xinxin Xu is shaping the next generation of intelligent, context-aware machines.
Research Focus
Key Achievements
Top Papers
- 1A survey of language-grounded multimodal 3D scene understanding5 citations · 2025