Papers
7
Total Citations
91
H-Index
5
About
Zhijun Fang is a leading researcher in 3D computer vision, with a focus on depth estimation, multi-modal perception, and 3D scene understanding for intelligent systems. His work bridges the gap between cost-effective monocular video analysis and high-precision 3D sensing, enabling robots and autonomous agents to perceive depth and ego-motion without expensive LiDAR sensors. Fang’s most cited paper, "Multi-modal 3D object detection by 2D-guided precision anchor proposal and multi-layer fusion" (2021, 38 citations), introduces a novel framework that leverages 2D image cues to guide 3D detection, significantly improving accuracy in cluttered environments. His 2018 paper on unsupervised depth estimation using perceptual losses (24 citations) pioneered a learning-based approach to recover 3D information from monocular video, a foundational contribution to self-supervised depth learning. More recently, his work on self-supervised depth and ego-motion for human-computer interaction (2023, 10 citations) and efficient RGB-D fusion for 6D pose estimation (2022, 9 citations) demonstrates his ongoing impact in making 3D perception accessible and robust. With over 90 total citations and a trajectory toward lightweight, real-time models like LiDUT-Depth (2024), Fang is shaping the future of affordable 3D vision for robotics, autonomous driving, and interactive systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2Depth Estimation of Video Sequences With Perceptual Losses24 citations · 2018
- 3
- 4EFN6D: an efficient RGB-D fusion network for 6D pose estimation9 citations · 2022
- 5
- 6
- 7Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation1 citations · 2025