Home /Research /Multi-modal LLM-enabled Long-horizon Skill Learning for Robotic Manipulation
MANIPULATION

Multi-modal LLM-enabled Long-horizon Skill Learning for Robotic Manipulation

Runjia Tan, Shanhe Lou, Yanxin Zhou, Chen Lv

Year
2024
Citations
4

Abstract

The advent of Large Language Models (LLMs) has empowered robots to execute tasks based on human instructions. Nonetheless, the challenge still persists in endowing robots with the capability to learn from the progress during interacting with human, which blocks the application of robots on human-like assistant. In response, this paper proposes a LLMs-based framework aimed at facilitating robots in acquiring new skills through interaction history. The proposed framework comprises three integral components: 1) a subsystem named the Fast Learner, consisting of three functionally distinct GPT-based modules. These modules are adept at decomposing long-horizon instructions into a sequence of manageable tasks while synthesizing skill schemas from historical interactions. 2) a hierarchical skill library categorizing tasks based on complexity, alongside a collection of meta-tasks that encompass all other tasks. 3) a scene understanding module which identifies regions of interest and generates relationship graphs based on visual input combined with textual prompts. Comprehensive evaluation of the system is conducted in both simulated environments and with real robots, demonstrating its exceptional proficiency in long-horizon task decomposition and effectiveness in acquiring and mastering new skills.

Keywords

ModalComputer scienceHorizonArtificial intelligenceHuman–computer interactionMathematics

Related papers

Browse all MANIPULATION papers