Robot Arm Grasping based on Multi-threaded PPO Reinforcement Learning Algorithm
Aiqing Bi
- Year
- 2024
- Citations
- 3
Abstract
In recent years, the problem of robotic arm grasping has been the research focus in the field of robotic arm. This thesis carries out research on robotic arm grasping based on the multi-threaded PPO (proximal policy optimization) algorithm, and establishes a robotic arm grasping system model by using the multi-threaded PPO algorithm. The original PPO algorithm is difficult to select the appropriate policy, and the difference between the old and the new policy varies a lot during the training process, which affects the training effect, whereas the current PPO utilizes the CLIP function to reduce the gradient calculation process. Therefore, the reward function is designed based on the current PPO algorithm, and the method of multi-threaded parallel computing, dominant value regularization and reward scaling is proposed, which further improves the training stability and scalability of PPO, and accelerates the sampling efficiency and convergence of the algorithm.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002