Home /Research /Generating Robot Action Sequences: An Efficient Vision-Language Models with Visual Prompts
OTHER

Generating Robot Action Sequences: An Efficient Vision-Language Models with Visual Prompts

Weihao CAI, Yoshiki Mori, Nobutaka Shimada

Year
2024
Citations
2

Abstract

This paper focuses on equipping robots with the capability to comprehend human commands, plan appropriate actions, and execute them effectively. Recent research advances suggest that Large Language Models (LLMs) with Visual Language Model (VLM) components show promising prospects in robot task planning. However, they sometimes encounter difficulties when dealing with ambiguous scenarios, such as understanding spatial relationships between objects or distinguishing between different degrees of drawer openings. This paper presents a novel approach for generating robot-executable action sequences using visual prompts and Vision-Language Models (VLMs). Our approach enhances the robot's environmental understanding by using a set of annotations (boxes, labels) as input. It generates action sequences based on the current environment to achieve goals and creates new sequences in case of action failure or external disturbances. Experimental results demonstrate outstanding performance in executing complex tasks, significantly improving the robot's task completion capabilities. The proposed technique offers a promising solution for more efficient and adaptable robot behavior acquisition.

Keywords

Computer scienceAction (physics)Artificial intelligenceComputer visionRobotHuman–computer interaction

Related papers

Browse all OTHER papers