Learning Complicated Manipulation Skills Via Deterministic Policy with Limited Demonstrations
Haofeng Liu, Jiayi Tan, Yiwen Chen, Marcelo H. Ang
- 发表年份
- 2023
- 引用次数
- 2
摘要
Deep reinforcement learning, when combined with demonstrations, can effectively formulate policies for manipulators. However, the practical collection of ample high-quality demonstrations is time-consuming, and demonstrations generated by humans may not perfectly correspond with the operational demands of robots. These challenges are intensified by issues such as non-Markovian processes and excessive reliance on demonstrations. Our study indicates that in manipulation tasks, reinforcement learning (RL) agents are sensitive to the quality of demonstrations and struggle to adapt to those derived from humans. As a result, leveraging low-quality or scarce demonstrations to assist reinforcement learning in developing superior policies presents a significant challenge. In some cases, dependence on limited demonstrations may paradoxically impair performance. To address these challenges, we propose a novel algorithm, TD3fG (TD3 learning from a generator). This algorithm facilitates a seamless transition from learning from experts to learning from experience, enabling agents to assimilate prior knowledge while mitigating the negative impacts of the demonstrations. Our algorithm demonstrates notable improvement in the Adroit manipulator and MuJoCo tasks, even with limited demonstrations of mixed failure trajectory.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002