Multi-modal LLM-enabled Long-horizon Skill Learning for Robotic Manipulation
Runjia Tan, Shanhe Lou, Yanxin Zhou, Chen Lv
- Year
- 2024
- Citations
- 4
Abstract
The advent of Large Language Models (LLMs) has empowered robots to execute tasks based on human instructions. Nonetheless, the challenge still persists in endowing robots with the capability to learn from the progress during interacting with human, which blocks the application of robots on human-like assistant. In response, this paper proposes a LLMs-based framework aimed at facilitating robots in acquiring new skills through interaction history. The proposed framework comprises three integral components: 1) a subsystem named the Fast Learner, consisting of three functionally distinct GPT-based modules. These modules are adept at decomposing long-horizon instructions into a sequence of manageable tasks while synthesizing skill schemas from historical interactions. 2) a hierarchical skill library categorizing tasks based on complexity, alongside a collection of meta-tasks that encompass all other tasks. 3) a scene understanding module which identifies regions of interest and generates relationship graphs based on visual input combined with textual prompts. Comprehensive evaluation of the system is conducted in both simulated environments and with real robots, demonstrating its exceptional proficiency in long-horizon task decomposition and effectiveness in acquiring and mastering new skills.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991