首页 /研究 /Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks
HRI

Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks

Michelle A. Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Silvio Savarese, Li Fei-Fei, Animesh Garg, Jeannette Bohg

发表年份
2019
引用次数
2
访问权限
开放获取

摘要

Contact-rich manipulation tasks in unstructured environments often require\nboth haptic and visual feedback. It is non-trivial to manually design a robot\ncontroller that combines these modalities which have very different\ncharacteristics. While deep reinforcement learning has shown success in\nlearning control policies for high-dimensional inputs, these algorithms are\ngenerally intractable to deploy on real robots due to sample complexity. In\nthis work, we use self-supervision to learn a compact and multimodal\nrepresentation of our sensory inputs, which can then be used to improve the\nsample efficiency of our policy learning. Evaluating our method on a peg\ninsertion task, we show that it generalizes over varying geometries,\nconfigurations, and clearances, while being robust to external perturbations.\nWe also systematically study different self-supervised learning objectives and\nrepresentation learning architectures. Results are presented in simulation and\non a physical robot.\n

关键词

Reinforcement learningComputer scienceHaptic technologyRepresentation (politics)Artificial intelligenceModalitiesTask (project management)RobotHuman–computer interactionController (irrigation)

相关论文

查看 HRI 分类全部论文