Grasping Control of a Vision Robot Based on a Deep Attentive Deterministic Policy Gradient
Xiangxin Ji, Feng Xiong, Weichang Kong, Dongfei Wei, Zeyan Shen
- 发表年份
- 2021
- 引用次数
- 4
- 访问权限
- 开放获取
摘要
Reinforcement learning can achieve excellent performance in the field of robotic grasping if the grasping target is stable. However, during applications in the real world, robot needs to overcome the effects of a complex working environment with different types of target objects, so it is more difficult to maintain the quality of action planning, even in the same scene. In order to make an agent have the ability to plan actions in a more adaptive way, the deep attentive deterministic policy gradient algorithm is applied in this article. An attention region proposal network is used to select the message of the pre-exploration area. Then this message is calculated using the adaptive exploration method to regulate the strategy as the target changes. Furthermore, a stratified reward function, which is used to reduce the negative influence of miscellaneous information brought by the sparse reward matrix, is defined according to the distance between the end effector and the center of the pre-exploration area. The results show that the DADPG is able to produce a robust strategy with noise interference, and can train in a more efficient way due to the hierarchical reward function.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002