6D Pose Estimation for Vision-guided Robot Grasping Based on Monocular Camera
Shiman Wang, Jianxin Liu, Qiang Lu, Zhen Liu, Yongjie Zeng, Dechang Zhang, Bo Chen
- Year
- 2023
- Citations
- 3
Abstract
Enhancing the intelligence of grasping operations through visual empowerment is currently a hot research topic in the field of robotics, where pose estimation is an important component. However, existing methods that solely rely on RGB images mostly employ direct methods or indirect methods using two-stage/multistage network predictions. Estimating the translation and rotation components of pose solely through a single network model does not yield optimal solutions. In this paper, different networks were used to predict the translation and rotation components, and convolutional block attention module (CBAM) and Pyramid Pooling Module (PPM) were introduced to address the issues that features decay in the backbone residual network after multiple layers of convolution and the insufficient utilization of feature information. Furthermore, in the rotation prediction branch, the problem of inaccurate recognition caused by weakened details after multiple layers of convolution was addressed by incorporating pyramid pooling into the network to capture richer detail and contextual information. Experimental tests were conducted on the publicly available Linemod dataset and in real environments. The results demonstrate that the proposed network exhibits a significant improvement of 2.76% in the most common ADD metric compared to a coordinate-based disentangled pose network (CDPN). It also achieves a 1 % improvement in the 5°, 5cm metrics and comparable performance in the Proj.2D metric with a 0.2% increase.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002