首页 /研究 /Watch and Act: Learning Robotic Manipulation From Visual Demonstration
MANIPULATION

Watch and Act: Learning Robotic Manipulation From Visual Demonstration

Shuo Yang, Wei Zhang, Ran Song, Jiyu Cheng, Hesheng Wang, Yibin Li

发表年份
2023
引用次数
31

摘要

Learning from demonstration holds the promise of enabling robots to learn diverse actions from expert experience. In contrast to learning from observation-action pairs, humans learn to imitate in a more flexible and efficient manner: learning behaviors by simply “watching.” In this article, we propose a “watch-and-act” imitation learning pipeline that endows a robot with the ability of learning diverse manipulations from visual demonstrations. Specifically, we address this problem by intuitively casting it as two subtasks: 1) understanding the demonstration video and 2) learning the demonstrated manipulations. First, a captioning module based on visual change is presented to understand the demonstration by translating the demonstration video into a command sentence. Then, to execute the captioning command, a manipulation module that learns the demonstrated manipulations is built upon an instance segmentation model and a manipulation affordance prediction model. We validate the superiority of the two modules over existing methods separately via extensive experiments and demonstrate the whole robotic imitation system developed based on the two modules in diverse scenarios using a real robotic arm. Supplementary video is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://vsislab.github.io/watch-and-act/</uri> .

关键词

Computer scienceHuman–computer interactionArtificial intelligenceComputer vision

相关论文

查看 MANIPULATION 分类全部论文