Qihang Zhang
Papers
2
Total Citations
28
H-Index
2
About
Qihang Zhang is a rising researcher at the intersection of computer vision, robotics, and autonomous driving, with a focus on enabling machines to perceive and interact with the physical world. His work bridges two critical frontiers: learning driving behaviors from unstructured video data and achieving robust 3D understanding of everyday objects. In his highly cited paper "Learning to Drive by Watching YouTube Videos" (18 citations), Zhang pioneered an action-conditioned contrastive policy pretraining method that allows autonomous agents to learn driving policies directly from human demonstrations in naturalistic settings, bypassing the need for expensive simulation environments. His second notable contribution, "Generative Category-Level Shape and Pose Estimation with Semantic Primitives" (10 citations), tackles the grand challenge of robotic 3D perception by proposing a novel framework that uses semantic primitives to handle the vast diversity of object shapes in unknown environments. This work represents a significant step toward empowering autonomous agents with the robust spatial understanding needed for real-world manipulation tasks. Zhang’s research exemplifies the growing trend of learning from diverse, unstructured data sources—whether YouTube videos or everyday objects—to build more adaptable and intelligent robotic systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2Generative Category-Level Shape and Pose Estimation with Semantic Primitives10 citations · 2022