首页 /研究 /Multi-modal LLM-enabled Long-horizon Skill Learning for Robotic Manipulation
MANIPULATION

Multi-modal LLM-enabled Long-horizon Skill Learning for Robotic Manipulation

Runjia Tan, Shanhe Lou, Yanxin Zhou, Chen Lv

发表年份
2024
引用次数
4

摘要

The advent of Large Language Models (LLMs) has empowered robots to execute tasks based on human instructions. Nonetheless, the challenge still persists in endowing robots with the capability to learn from the progress during interacting with human, which blocks the application of robots on human-like assistant. In response, this paper proposes a LLMs-based framework aimed at facilitating robots in acquiring new skills through interaction history. The proposed framework comprises three integral components: 1) a subsystem named the Fast Learner, consisting of three functionally distinct GPT-based modules. These modules are adept at decomposing long-horizon instructions into a sequence of manageable tasks while synthesizing skill schemas from historical interactions. 2) a hierarchical skill library categorizing tasks based on complexity, alongside a collection of meta-tasks that encompass all other tasks. 3) a scene understanding module which identifies regions of interest and generates relationship graphs based on visual input combined with textual prompts. Comprehensive evaluation of the system is conducted in both simulated environments and with real robots, demonstrating its exceptional proficiency in long-horizon task decomposition and effectiveness in acquiring and mastering new skills.

关键词

ModalComputer scienceHorizonArtificial intelligenceHuman–computer interactionMathematics

相关论文

查看 MANIPULATION 分类全部论文