首页 /研究 /K-VIL: Keypoints-Based Visual Imitation Learning
MANIPULATION

K-VIL: Keypoints-Based Visual Imitation Learning

Jianfeng Gao, Zhi Tao, Noémie Jaquier, Tamim Asfour

发表年份
2023
引用次数
24

摘要

Visual imitation learning provides efficient and intuitive solutions for robotic systems to acquire novel manipulation skills. However, simultaneously learning geometric task constraints and control policies from visual inputs alone remains a challenging problem. In this article, we propose the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">keypoint-based visual imitation learning</i> (K-VIL) approach that automatically extracts sparse, object-centric, and embodiment-independent task representations from a small number of human demonstration videos. The task representation is composed of keypoint-based geometric constraints on principal manifolds, their associated local frames, and the movement primitives that are then needed for the task execution. Our approach is capable of extracting such task representations from a single-demonstration video and of incrementally updating them when new demonstrations are available. To reproduce manipulation skills using the learned set of prioritized geometric constraints in novel scenes, we introduce a novel keypoint-based admittance controller. We evaluate our approach in several real-world applications, showcasing its ability to deal with cluttered scenes, viewpoint mismatch, new instances of categorical objects, and large object pose and shape variations. Our evaluation demonstrates the efficiency and robustness of our approach in both one-shot and few-shot imitation learning settings.

关键词

Computer scienceArtificial intelligenceComputer visionRobustness (evolution)Object (grammar)Task (project management)Representation (politics)

相关论文

查看 MANIPULATION 分类全部论文