首页 /研究 /Contextual Latent-Movements Off-Policy Optimization for Robotic Manipulation Skills
MANIPULATION

Contextual Latent-Movements Off-Policy Optimization for Robotic Manipulation Skills

Samuele Tosatto, Georgia Chalvatzaki, Jan Peters

发表年份
2020
访问权限
开放获取

摘要

Parameterized movement primitives have been extensively used for imitation learning of robotic tasks. However, the high-dimensionality of the parameter space hinders the improvement of such primitives in the reinforcement learning (RL) setting, especially for learning with physical robots. In this paper we propose a novel view on handling the demonstrated trajectories for acquiring low-dimensional, non-linear latent dynamics, using mixtures of probabilistic principal component analyzers (MPPCA) on the movements' parameter space. Moreover, we introduce a new contextual off-policy RL algorithm, named LAtent-Movements Policy Optimization (LAMPO). LAMPO can provide gradient estimates from previous experience using self-normalized importance sampling, hence, making full use of samples collected in previous learning iterations. These advantages combined provide a complete framework for sample-efficient off-policy optimization of movement primitives for robot learning of high-dimensional manipulation skills. Our experimental results conducted both in simulation and on a real robot show that LAMPO provides sample-efficient policies against common approaches in literature.

关键词

cs.ROcs.LG

相关论文

查看 MANIPULATION 分类全部论文