Home /Research /Goal-Conditioned Reinforcement Learning With Adaptive Intrinsic Curiosity and Universal Value Network Fitting for Robotic Manipulation
MANIPULATION

Goal-Conditioned Reinforcement Learning With Adaptive Intrinsic Curiosity and Universal Value Network Fitting for Robotic Manipulation

Zihao Sun, Xianfeng Yuan, Qingyang Xu, Bao Pang, Yong Song, Rui Song, Yibin Li

Year
2024
Citations
3

Abstract

Hindsight experience replay (HER) has greatly increased the possibility of using deep reinforcement learning (DRL) for robotic manipulation with sparse rewards. However, there are still concerns about low learning efficiency and poor performance due to its insufficient exploration ability and bias against the initial goal introduced by HER. In this article, to solve this problem, a multigoal robotic manipulation DRL method based on adaptive intrinsic curiosity and universal value network fitting (AIC-UVNF) is proposed to further improve the exploration ability and learning performance. Specifically, this method utilizes an improved curiosity mechanism to construct a joint intrinsic reward and adaptively adjust the proportion, which can enhance exploration ability and avoid excessive pursuit of novel states. In addition, a universal value network fitting approach is proposed to incorporate the initial goal into the value function fitting process, which employs the value of the initial goal to eliminate the bias of HER in the algorithm update. Combined with the off-policy soft actor-critic method, AIC-UVNF is verified on multigoal robotic manipulation tasks. The results show that the proposed method achieves better convergence efficiency and learning performance.

Keywords

Reinforcement learningCuriosityComputer scienceArtificial intelligenceValue (mathematics)ReinforcementArtificial neural networkControl theory (sociology)Control engineeringMachine learning

Related papers

Browse all MANIPULATION papers