Home /Research /Fitted Q-iteration by Advantage Weighted Regression
OTHER

Fitted Q-iteration by Advantage Weighted Regression

Gerhard Neumann, Jan Peters

Year
2008
Citations
40
Access
Open access

Abstract

Recently, fitted Q-iteration (FQI) based methods have become more popular due
\nto their increased sample efficiency, a more stable learning process and the higher
\nquality of the resulting policy. However, these methods remain hard to use for continuous
\naction spaces which frequently occur in real-world tasks, e.g., in robotics
\nand other technical applications. The greedy action selection commonly used for
\nthe policy improvement step is particularly problematic as it is expensive for continuous
\nactions, can cause an unstable learning process, introduces an optimization
\nbias and results in highly non-smooth policies unsuitable for real-world systems.
\nIn this paper, we show that by using a soft-greedy action selection the policy
\nimprovement step used in FQI can be simplified to an inexpensive advantage weighted
\nregression. With this result, we are able to derive a new, computationally
\nefficient FQI algorithm which can even deal with high dimensional action spaces.

Keywords

Computer scienceAction (physics)Process (computing)Quality (philosophy)Artificial intelligenceSelection (genetic algorithm)Action selectionGreedy algorithmMachine learningMathematical optimization

Related papers

Browse all OTHER papers