Multi-view dreaming: multi-view world model with contrastive learning
Akira Kinose, Masashi Okada, Ryo Okumura, Tadahiro Taniguchi
- Year
- 2023
- Citations
- 3
Abstract
AbstractIn this paper, we propose Multi-View Dreaming, a novel reinforcement learning agent for integrated recognition and control from multi-view observations by extending Dreaming. Most current reinforcement learning method assumes a single-view observation space, and this imposes limitations on the observed data, such as lack of spatial information and occlusions. This makes obtaining ideal observational information from the environment difficult and is a bottleneck for real-world robotics applications. In this paper, we use contrastive learning to train a shared latent space between different viewpoints and show how the Products of Experts approach can be used to integrate and control the probability distributions of latent states for multiple viewpoints. We also propose Multi-View DreamingV2, a variant of Multi-View Dreaming that uses a categorical distribution to model the latent state instead of the Gaussian distribution. Experiments show that the proposed method outperforms simple extensions of existing methods in a realistic robot control task.KEYWORDS: World modelsreinforcement learningmultimodalrobotic manipulationsensor integration Disclosure statementNo potential conflict of interest was reported by the author(s).Correction StatementThis article has been corrected with minor changes. These changes do not impact the academic content of the article.Additional informationNotes on contributorsAkira KinoseAkira Kinose is currently a research engineer at Panasonic Connect Co., Ltd. He received his M.Eng. in Information Science and Engineering from Ritsumeikan University in 2020. His research involves state representation learning and large-scale language models.Masashi OkadaMasashi Okada is currently a senior engineer at Panasonic Holdings Corp. He received his Ph.D. in Information Science from Osaka University in 2013. His research involves learning-based control methods and probabilistic machine.Ryo OkumuraRyo Okumura is currently a senior engineer at Panasonic Holdings Corp. He received his M.Eng. in Information Science and Technology from The University of Tokyo in 2009. From 2016 to 2018, he was a visiting scholar at Stanford University in the Department of Mechanical Engineering in the School of Engineering. His research involves state representation learning and robot control.Tadahiro TaniguchiTadahiro Taniguchi received his Ph.D. degree from Kyoto University, in 2006. From 2005 to 2008, he was a research fellow of the Japan Society for the Promotion of Science. From 2008 to 2010, he was an assistant professor at the Department of Human and Computer Intelligence, Ritsumeikan University. From 2010 to 2017, he was an associate professor in the same department. From 2015 to 2016, he was a visiting associate professor at the Department of Electrical and Electronic Engineering, Imperial College London. Since 2017, he has been a professor at the Department of Information Science and Engineering, Ritsumeikan University, and a visiting general chief scientist at Panasonic (Holdings) Corporation. He has been engaged in research on machine learning, emergent systems, and symbol emergence in robotics.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002