Shengyi Qian
Papers
5
Total Citations
78
H-Index
3
About
Shengyi Qian is a researcher at the forefront of 3D scene understanding, embodied AI, and robotic perception, with a focus on bridging language, vision, and spatial reasoning for real-world robotic applications. His most celebrated contribution, LLM-Grounder, introduced a novel framework leveraging large language models as agents for open-vocabulary 3D visual grounding, enabling robots to interpret complex natural language queries and locate objects in 3D environments without heavy reliance on labeled data — a paper that has already garnered over 60 citations since its publication. Building on this foundation, Qian contributed to 3D-GRAND, a million-scale dataset designed to reduce hallucination and improve grounding in 3D-aware large language models, addressing one of the field's most pressing reliability challenges. His work on understanding 3D object interaction from single images demonstrates a broader ambition to equip machines with human-like spatial intuition, while 3D-MVP advances robotic manipulation through multiview 3D pretraining. Collectively, Qian's research tackles the critical gap between language understanding and physical world perception, making meaningful strides toward intelligent, spatially-aware robotic systems capable of operating in complex, unstructured environments.
Research Focus
Key Achievements
Top Papers
- 1
- 2Understanding 3D Object Interaction from a Single Image8 citations · 2023
- 3
- 43D-MVP: 3D Multiview Pretraining for Manipulation3 citations · 2025
- 5