A reinforcement learning accelerated by state space reduction
Kei Senda, S. Mano, Shinji Fujii
- 发表年份
- 2003
- 引用次数
- 4
摘要
This paper discusses a method to accelerate reinforcement learning. A concept is firstly defined, i.e., the state space reduction conserving policy. An algorithm is then given, where the optimal cost-to-go and the optimal policy of the reduced space is calculated from those of the original space. Using the reduced state space, learning convergence is accelerated. Its usefulness for DP iteration and Q-learning are compared through a maze example. The convergence of the optimal cost-to-go in the original state space needs approximately N or more times as long as that in the reduced state space, where N is a ratio of the number of the original states to the reduced. The acceleration effect for Q-learning is more remarkable than that for the DP iteration. The proposal technique is also applied to a robot manipulator working for a peg-in-hole task with geometric constraints. The state space reduction can be considered as a model of the change of observation, i.e., one of cognitive actions. The obtained results explain that the change of observation is reasonable in terms of learning efficiency.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002