Haoyu Zhen
Papers
1
Total Citations
14
H-Index
1
About
Haoyu Zhen is a rising researcher at the forefront of embodied AI and 3D vision-language reasoning. His work focuses on bridging the gap between high-level semantic understanding and low-level robotic control, with a particular emphasis on integrating 3D spatial awareness into generative world models. His most notable contribution, the 2024 paper "3D-VLA: A 3D Vision-Language-Action Generative World Model," addresses a critical limitation in existing vision-language-action (VLA) models, which typically operate on 2D inputs and learn direct perception-to-action mappings. Zhen's 3D-VLA framework instead introduces a generative world model that reasons about 3D physical dynamics and object relations, enabling more robust and generalizable robotic manipulation. Though early in his career, this work has already garnered 14 citations, signaling strong interest from the community. By pioneering the fusion of 3D scene understanding with action generation, Zhen is laying the groundwork for robots that can truly comprehend and interact with the three-dimensional world, making him a promising voice in the next wave of embodied intelligence research.
Research Focus
Key Achievements
Top Papers
- 13D-VLA: A 3D Vision-Language-Action Generative World Model14 citations · 2024