Beichen Wang

New York University

Papers

2

Total Citations

13

H-Index

2

About

Beichen Wang is forging a bold new path at the intersection of computer vision, natural language processing, and embodied AI. His primary research focuses on leveraging Vision Language Models (VLMs) to bridge the gap between human demonstration and robotic execution. Wang’s most significant contribution is the development of a novel framework that enables robots to interpret unconstrained human demonstration videos and translate them directly into executable action plans. This work, detailed in his highly cited paper “VLM See, Robot Do,” tackles the critical challenge of generalizability in robotics by harnessing VLMs’ powerful common sense reasoning. By moving beyond language-only instructions, his research allows robots to learn from rich, visual human examples, dramatically simplifying the robot programming pipeline. With his first-author work already accumulating over a dozen citations shortly after publication, Wang is establishing himself as a rising star in robot learning. His achievements demonstrate a rare ability to synthesize complex multimodal models into practical, scalable solutions, positioning him at the forefront of creating robots that can truly understand and replicate human behavior from visual observation alone.

Research Focus

Key Achievements

2
H-Index
2
Papers
13
Total Citations
7
Avg Citations/Paper
🏆 Most Cited Paper
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
10 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 3
🏛 Institutions: New York University

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago