Fitted Q-iteration by Advantage Weighted Regression
Gerhard Neumann, Jan Peters
- Year
- 2008
- Citations
- 40
- Access
- Open access
Abstract
Recently, fitted Q-iteration (FQI) based methods have become more popular due \nto their increased sample efficiency, a more stable learning process and the higher \nquality of the resulting policy. However, these methods remain hard to use for continuous \naction spaces which frequently occur in real-world tasks, e.g., in robotics \nand other technical applications. The greedy action selection commonly used for \nthe policy improvement step is particularly problematic as it is expensive for continuous \nactions, can cause an unstable learning process, introduces an optimization \nbias and results in highly non-smooth policies unsuitable for real-world systems. \nIn this paper, we show that by using a soft-greedy action selection the policy \nimprovement step used in FQI can be simplified to an inexpensive advantage weighted \nregression. With this result, we are able to derive a new, computationally \nefficient FQI algorithm which can even deal with high dimensional action spaces.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991