Papers

3

Total Citations

51

H-Index

3

About

Zhengkai Jiang is at the forefront of integrating large language models (LLMs) and vision-language models (VLMs) with robotic manipulation, pioneering frameworks that bridge high-level language understanding with low-level physical action. His key research areas include multi-modal instruction following, robotic affordance learning, and 3D object-centric manipulation. Jiang’s major contribution is the development of **Instruct2Act** (30 citations), a seminal framework that maps multi-modal instructions—text, images, and demonstrations—into sequential robotic actions using LLMs, enabling robots to interpret complex human commands. He further advanced the field with **ManipVQA** (18 citations), which injects robotic affordance and physically grounded knowledge into multi-modal LLMs, significantly improving manipulation task performance by addressing the critical gap in robotics-specific reasoning. His most recent work, **UniAff** (2025), introduces a unified representation of affordances for tool usage and articulation, integrating 3D motion constraints with task understanding. Jiang’s research is highly impactful, with his papers quickly accumulating citations and shaping the next generation of intelligent robotic systems. His work stands out for its practical focus on enabling robots to understand not just *what* to do, but *how* to physically interact with the world.

Research Focus

Key Achievements

3
H-Index
3
Papers
51
Total Citations
17
Avg Citations/Paper
🏆 Most Cited Paper
Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
30 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 18
🏛 Institutions: University College of Applied Science, Hong Kong University of Science and Technology

Top Papers

  1. 1
  2. 2
  3. 3

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago