Papers
3
Total Citations
51
H-Index
3
About
Zhengkai Jiang is at the forefront of integrating large language models (LLMs) and vision-language models (VLMs) with robotic manipulation, pioneering frameworks that bridge high-level language understanding with low-level physical action. His key research areas include multi-modal instruction following, robotic affordance learning, and 3D object-centric manipulation. Jiang’s major contribution is the development of **Instruct2Act** (30 citations), a seminal framework that maps multi-modal instructions—text, images, and demonstrations—into sequential robotic actions using LLMs, enabling robots to interpret complex human commands. He further advanced the field with **ManipVQA** (18 citations), which injects robotic affordance and physically grounded knowledge into multi-modal LLMs, significantly improving manipulation task performance by addressing the critical gap in robotics-specific reasoning. His most recent work, **UniAff** (2025), introduces a unified representation of affordances for tool usage and articulation, integrating 3D motion constraints with task understanding. Jiang’s research is highly impactful, with his papers quickly accumulating citations and shaping the next generation of intelligent robotic systems. His work stands out for its practical focus on enabling robots to understand not just *what* to do, but *how* to physically interact with the world.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3