Haoyu Zhen

Papers

1

Total Citations

14

H-Index

1

About

Haoyu Zhen is a rising researcher at the forefront of embodied AI and 3D vision-language reasoning. His work focuses on bridging the gap between high-level semantic understanding and low-level robotic control, with a particular emphasis on integrating 3D spatial awareness into generative world models. His most notable contribution, the 2024 paper "3D-VLA: A 3D Vision-Language-Action Generative World Model," addresses a critical limitation in existing vision-language-action (VLA) models, which typically operate on 2D inputs and learn direct perception-to-action mappings. Zhen's 3D-VLA framework instead introduces a generative world model that reasons about 3D physical dynamics and object relations, enabling more robust and generalizable robotic manipulation. Though early in his career, this work has already garnered 14 citations, signaling strong interest from the community. By pioneering the fusion of 3D scene understanding with action generation, Zhen is laying the groundwork for robots that can truly comprehend and interact with the three-dimensional world, making him a promising voice in the next wave of embodied intelligence research.

Research Focus

Key Achievements

1
H-Index
1
Papers
14
Total Citations
14
Avg Citations/Paper
🏆 Most Cited Paper
3D-VLA: A 3D Vision-Language-Action Generative World Model
14 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 7

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago