3D Space Perception via Disparity Learning Using Stereo Images and an Attention Mechanism: Real-Time Grasping Motion Generation for Transparent Objects
Xianqiao Cai, Hiroshi Ito, Hyogo Hiruma, Tetsuya Ogata
- 发表年份
- 2024
- 引用次数
- 2
摘要
Object grasping in 3D space is crucial for robotic applications. Such tasks are performed by utilizing depth map data acquired from RGB-D images or 3D point cloud data. However, these methods struggle when dealing with transparent objects, as transparency limits sensor performance when predicting depth maps. Additionally, the grasping motions are predicted without incorporating the relationship between depth data and motion information, which limits the motion's flexibility. In this letter, to address these problems, we propose an end-to-end motion generation model using stereo RGB images, a deep-learning model that incorporates image and motion information. Furthermore, visual attention mechanisms are used for extracting task-related attention points, which is essential for building spatial cognition constructs. Real-robot experimental results confirmed that the proposed model is able to grasp transparent objects under various situations, including unseen positions, heights, and background. It was also found that the model self-organized a spatial cognition representation within its hidden states, suggesting that the integrated learning of robot motion and stable spatial attention points is important for spatial perception. Such explicit feature representations cannot be obtained via learning motion alone.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002