Shuwen Dong
Papers
2
Total Citations
13
H-Index
2
About
Shuwen Dong is a rising researcher at the forefront of embodied AI and robot learning, with a focus on bridging human demonstration and robotic action through vision-language models (VLMs). In their highly influential work, "VLM See, Robot Do," Dong pioneers a novel framework that enables robots to interpret unlabeled human demonstration videos and directly translate them into executable action plans. This approach leverages the common-sense reasoning and generalization power of large VLMs, bypassing the traditional need for explicit language instructions or extensive robot-specific training data. The paper has rapidly garnered over 10 citations within its first year, signaling its immediate impact on the robotics community. Dong’s key contribution lies in demonstrating how VLMs can serve as a versatile interface for robot learning, effectively allowing robots to "see" and "do" by understanding human intent from raw video. This work not only advances task and motion planning but also opens new pathways for scalable robot learning from human demonstrations. As a young researcher, Dong is already shaping the next generation of intuitive, human-centric robotic systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2