Home /Research /Talk Through It: End User Directed Manipulation Learning
MANIPULATION

Talk Through It: End User Directed Manipulation Learning

C. Winge, Adam Imdieke, Bahaa Aldeeb, Dongyeop Kang, Karthik Desingh

Year
2024
Citations
4

Abstract

Training robots to perform a huge range of tasks in many different environments is immensely difficult. Instead, we propose selectively training robots based on end-user preferences. Given a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">factory model</i> that lets an end user instruct a robot to perform lower-level actions (e.g. ‘Move left’), we show that end users can collect demonstrations using language to train their <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">home model</i> for higher-level tasks specific to their needs (e.g. ‘Open the top drawer and put the block inside’). We demonstrate this framework on robot manipulation tasks using RLBench environments. Our method results in a 13% improvement in task success rates compared to a baseline method. We also explore the use of the large vision-language model (VLM), Bard, to automatically break down tasks into sequences of lower-level instructions, aiming to bypass end-user involvement. The VLM is unable to break tasks down to our lowest level, but does achieve good results breaking high-level tasks into mid-level skills.

Keywords

Computer scienceHuman–computer interactionEnd userMultimediaWorld Wide Web

Related papers

Browse all MANIPULATION papers