Papers
8
Total Citations
149
H-Index
4
About
Cunjun Yu is a robotics researcher whose work lies at the intersection of computer vision, natural language processing, and robot manipulation. Their key contributions span visual servoing, interactive grasping, and socially-aware navigation. Yu’s most influential work, “Siamese Convolutional Neural Network for Sub-millimeter-accurate Camera Pose Estimation and Visual Servoing” (64 citations), introduced a novel Siamese architecture that achieves exceptional precision in guiding robot motions from camera images—a critical capability for high-accuracy tasks. Building on this, the INVIGORATE system (47 citations) pioneered interactive visual grounding and grasping in cluttered environments, enabling robots to understand natural language commands and retrieve specific objects even when occluded or stacked. More recently, Yu’s work on vision-language foundation models (2023) demonstrated how fine-tuning large multimodal models can produce effective robot imitators, bridging the gap between language understanding and physical action. Their latest research introduces GSON, a group-based social navigation framework leveraging large multimodal models for socially-aware robot movement, and Robi Butler, a household robot assistant enabling remote multimodal interaction. With over 140 total citations across these works, Yu is advancing the frontier of robots that can see, understand, and act in human environments.
Research Focus
Key Achievements
Top Papers
- 1
- 2INVIGORATE: Interactive Visual Grounding and Grasping in Clutter47 citations · 2021
- 3Vision-Language Foundation Models as Effective Robot Imitators19 citations · 2023
- 4
- 5INVIGORATE: Interactive Visual Grounding and Grasping in Clutter4 citations · 2021
- 6
- 7
- 8