Juexiao Zhang

New York University

Papers

2

Total Citations

13

H-Index

2

About

Juexiao Zhang is a rising researcher at the intersection of robotics, computer vision, and artificial intelligence, with a focused interest in leveraging large vision-language models (VLMs) for robotic manipulation and planning. His primary contribution lies in pioneering a novel paradigm that bridges human demonstration and robot execution: enabling robots to directly interpret and act upon unstructured human demonstration videos. In his highly cited work, "VLM See, Robot Do," Zhang demonstrates how VLMs can be used to translate raw human video into actionable robot plans, bypassing the need for explicit programming or extensive simulation data. This approach leverages the strong common-sense reasoning and generalization capabilities of modern VLMs to understand human intent and scene dynamics. While his most prominent paper has garnered over 10 citations in a short period, signaling immediate impact in the field, his ongoing work continues to refine these methods for more robust and generalizable robot learning. Zhang’s research is particularly notable for its potential to democratize robotics, allowing non-experts to teach robots complex tasks simply by showing them a video.

Research Focus

Key Achievements

2
H-Index
2
Papers
13
Total Citations
7
Avg Citations/Paper
🏆 Most Cited Paper
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
10 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 3
🏛 Institutions: New York University

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago