首页 /研究 /Self-Supervised Policy Adaptation during Deployment
MANIPULATION

Self-Supervised Policy Adaptation during Deployment

Nicklas Hansen, Rishabh Jangir, Yu Sun, Guillem Alenyà, Pieter Abbeel, Alexei A. Efros, Lerrel Pinto, Xiaolong Wang

发表年份
2021
引用次数
8

摘要

In most real world scenarios, a policy trained by reinforcement learning in\none environment needs to be deployed in another, potentially quite different\nenvironment. However, generalization across different environments is known to\nbe hard. A natural solution would be to keep training after deployment in the\nnew environment, but this cannot be done if the new environment offers no\nreward signal. Our work explores the use of self-supervision to allow the\npolicy to continue training after deployment without using any rewards. While\nprevious methods explicitly anticipate changes in the new environment, we\nassume no prior knowledge of those changes yet still obtain significant\nimprovements. Empirical evaluations are performed on diverse simulation\nenvironments from DeepMind Control suite and ViZDoom, as well as real robotic\nmanipulation tasks in continuously changing environments, taking observations\nfrom an uncalibrated camera. Our method improves generalization in 31 out of 36\nenvironments across various tasks and outperforms domain randomization on a\nmajority of environments.

关键词

Software deploymentComputer scienceReinforcement learningSuiteGeneralizationAdaptation (eye)Artificial intelligenceMachine learningDomain (mathematical analysis)Human–computer interaction

相关论文

查看 MANIPULATION 分类全部论文