Yichen Zhu
Papers
1
Total Citations
5
H-Index
1
About
Yichen Zhu is an emerging researcher at the intersection of robotics, 3D perception, and foundation models, with a particular focus on advancing robotic manipulation through the integration of spatial understanding and vision-language intelligence. His most recognized work, *PointVLA: Injecting the 3D World Into Vision-Language-Action Models*, addresses a fundamental limitation in modern robotic AI: the over-reliance on 2D RGB inputs in Vision-Language-Action (VLA) models, which constrains spatial reasoning in real-world environments. Rather than pursuing computationally prohibitive full retraining, Zhu's approach offers an efficient pathway to inject 3D world representations into pretrained VLA architectures — a contribution that bridges the gap between large-scale 2D pretraining and the geometric demands of physical interaction. Though his publication record is still developing, with *PointVLA* accumulating 5 citations since 2026, his work tackles a timely and consequential challenge in embodied AI. Researchers working on robot learning, 3D scene understanding, or multimodal foundation models will find Zhu's contributions particularly relevant as the field moves toward more spatially-aware, deployable robotic systems.
Research Focus
Key Achievements
Top Papers
- 1PointVLA: Injecting the 3D World Into Vision-Language-Action Models5 citations · 2026