Generating Robot Action Sequences: An Efficient Vision-Language Models with Visual Prompts
Weihao CAI, Yoshiki Mori, Nobutaka Shimada
- Year
- 2024
- Citations
- 2
Abstract
This paper focuses on equipping robots with the capability to comprehend human commands, plan appropriate actions, and execute them effectively. Recent research advances suggest that Large Language Models (LLMs) with Visual Language Model (VLM) components show promising prospects in robot task planning. However, they sometimes encounter difficulties when dealing with ambiguous scenarios, such as understanding spatial relationships between objects or distinguishing between different degrees of drawer openings. This paper presents a novel approach for generating robot-executable action sequences using visual prompts and Vision-Language Models (VLMs). Our approach enhances the robot's environmental understanding by using a set of annotations (boxes, labels) as input. It generates action sequences based on the current environment to achieve goals and creates new sequences in case of action failure or external disturbances. Experimental results demonstrate outstanding performance in executing complex tasks, significantly improving the robot's task completion capabilities. The proposed technique offers a promising solution for more efficient and adaptable robot behavior acquisition.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991