Research on PPO Algorithm Based on Continuous Robot
Dongyang Zhang, Ying Liang, Feng Zhang, Long Cui
- Year
- 2024
- Citations
- 2
Abstract
In order to solve the problems of slow convergence speed of the traditional deep reinforcement learning proximal strategy optimization (PPO) algorithm, an RPPO algorithm for continuous robot gait control with improved PPO algorithm was proposed, which introduced layer normalization in the policy network and value network to improve the convergence speed in the training process, and added L2 regularization to the loss function of the strategy network to help the neural network model learn more effective strategies to achieve better reward values. The improved algorithm was simulated and verified in the Mujooc simulation environment under Openai Gym. The simulation results show that the improved algorithm is better than the traditional algorithm and has better performance.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002