Zhizheng Zhang

Papers

3

Total Citations

8

H-Index

2

About

Zhizheng Zhang is an emerging researcher at the intersection of embodied AI, robotics, and vision-language understanding. His work focuses on equipping intelligent agents with the ability to perceive, reason, and act within complex real-world environments by leveraging the power of large-scale Vision-Language Models (VLMs). A central theme across his research is bridging the gap between language-grounded instructions and physical execution, whether in robotic manipulation or autonomous navigation. His benchmark work, Open6DOR, addresses a critical gap in object manipulation research by introducing rigorous 6-DoF rearrangement evaluation, pushing the field toward more realistic and generalizable robotic systems. In parallel, his NaVid series tackles Vision-and-Language Navigation (VLN), with NaVid introducing video-based VLM planning to improve generalization across unseen environments and sim-to-real transfer, while the follow-up NaVid-4D advances spatial-temporal reasoning through egocentric RGB-D video understanding for greater action precision. Though early in his career, Zhang's publications have already accumulated citations across the community, signaling growing recognition. His research offers promising directions for developing robots and navigation agents capable of following open-ended human instructions with human-like spatial intelligence.

Research Focus

Key Achievements

2
H-Index
3
Papers
8
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Open6DOR: Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach
5 citations · 2024
📈 Most Prolific Year: 2024 (2 Papers)
🤝 Key Collaborators: 23

Top Papers

  1. 1
  2. 2
  3. 3

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 16 days ago