Irving Fang

New York University

Papers

4

Total Citations

21

H-Index

3

About

Irving Fang is a rising star at the intersection of computer vision, robotics, and embodied AI. His research focuses on enabling robots to perceive, reason, and act by bridging the gap between human demonstrations and robotic execution. Fang’s most impactful work, “VLM See, Robot Do,” introduces a novel framework that leverages Vision Language Models (VLMs) to translate human demonstration videos directly into actionable robot plans—bypassing the need for complex manual programming. This work has already garnered over 10 citations since its 2024 publication, signaling its influence in the robotics community. In “FusionSense,” Fang pioneers a multi-modal approach that integrates common-sense reasoning with vision and tactile feedback for robust 3D reconstruction from sparse data, a critical step toward more perceptive robots. His work on “EgoPAT3Dv2” advances human-robot interaction by predicting 3D action targets from egocentric video, enhancing robot anticipation and safety. Together, Fang’s contributions demonstrate a clear trajectory toward robots that learn from human behavior, fuse diverse sensory inputs, and act intelligently in the real world.

Research Focus

Key Achievements

3
H-Index
4
Papers
21
Total Citations
5
Avg Citations/Paper
🏆 Most Cited Paper
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
10 citations · 2024
📈 Most Prolific Year: 2024 (2 Papers)
🤝 Key Collaborators: 21
🏛 Institutions: New York University

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago