Shuwen Dong

New York University

Papers

2

Total Citations

13

H-Index

2

About

Shuwen Dong is a rising researcher at the forefront of embodied AI and robot learning, with a focus on bridging human demonstration and robotic action through vision-language models (VLMs). In their highly influential work, "VLM See, Robot Do," Dong pioneers a novel framework that enables robots to interpret unlabeled human demonstration videos and directly translate them into executable action plans. This approach leverages the common-sense reasoning and generalization power of large VLMs, bypassing the traditional need for explicit language instructions or extensive robot-specific training data. The paper has rapidly garnered over 10 citations within its first year, signaling its immediate impact on the robotics community. Dong’s key contribution lies in demonstrating how VLMs can serve as a versatile interface for robot learning, effectively allowing robots to "see" and "do" by understanding human intent from raw video. This work not only advances task and motion planning but also opens new pathways for scalable robot learning from human demonstrations. As a young researcher, Dong is already shaping the next generation of intuitive, human-centric robotic systems.

Research Focus

Key Achievements

2
H-Index
2
Papers
13
Total Citations
7
Avg Citations/Paper
🏆 Most Cited Paper
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
10 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 3
🏛 Institutions: New York University

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago