Zhuoyang Zhang
Papers
2
Total Citations
10
H-Index
2
About
Zhuoyang Zhang is a rising researcher in the intersection of computer vision, human-robot interaction, and multimodal perception. His primary research focuses on gaze-directed visual grounding, a technique that integrates human gaze patterns with natural language to help robots understand ambiguous or under-specified object references. Zhang’s most significant contribution is the Gaze-directed Visual Grounding Network (GVGNet), introduced in his 2023 paper, which addresses the challenge of disambiguating human referring intentions when language alone is insufficient. This work has garnered 8 citations and lays the foundation for more intuitive human-robot collaboration. Building on this, his 2024 study explores gaze-assisted visual grounding via knowledge distillation for referred object grasping, achieving 2 citations and extending the framework to physical robotic manipulation. By combining gaze tracking with referring expression comprehension and segmentation, Zhang’s research enables robots to interpret subtle human cues—such as where a person looks while saying “that cup”—making interactions more natural and efficient. His work is particularly impactful for assistive robotics, manufacturing, and autonomous systems, where precise object identification is critical. Zhang’s innovative integration of gaze and language signals marks him as a promising contributor to embodied AI and human-centered robotics.
Research Focus
Key Achievements
Top Papers
- 1
- 2