Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks
Michelle A. Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Silvio Savarese, Li Fei-Fei, Animesh Garg, Jeannette Bohg
- Year
- 2019
- Citations
- 2
- Access
- Open access
Abstract
Contact-rich manipulation tasks in unstructured environments often require\nboth haptic and visual feedback. It is non-trivial to manually design a robot\ncontroller that combines these modalities which have very different\ncharacteristics. While deep reinforcement learning has shown success in\nlearning control policies for high-dimensional inputs, these algorithms are\ngenerally intractable to deploy on real robots due to sample complexity. In\nthis work, we use self-supervision to learn a compact and multimodal\nrepresentation of our sensory inputs, which can then be used to improve the\nsample efficiency of our policy learning. Evaluating our method on a peg\ninsertion task, we show that it generalizes over varying geometries,\nconfigurations, and clearances, while being robust to external perturbations.\nWe also systematically study different self-supervised learning objectives and\nrepresentation learning architectures. Results are presented in simulation and\non a physical robot.\n
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002