首页 /研究 /Fitted Q-iteration by Advantage Weighted Regression
OTHER

Fitted Q-iteration by Advantage Weighted Regression

Gerhard Neumann, Jan Peters

发表年份
2008
引用次数
40
访问权限
开放获取

摘要

Recently, fitted Q-iteration (FQI) based methods have become more popular due
\nto their increased sample efficiency, a more stable learning process and the higher
\nquality of the resulting policy. However, these methods remain hard to use for continuous
\naction spaces which frequently occur in real-world tasks, e.g., in robotics
\nand other technical applications. The greedy action selection commonly used for
\nthe policy improvement step is particularly problematic as it is expensive for continuous
\nactions, can cause an unstable learning process, introduces an optimization
\nbias and results in highly non-smooth policies unsuitable for real-world systems.
\nIn this paper, we show that by using a soft-greedy action selection the policy
\nimprovement step used in FQI can be simplified to an inexpensive advantage weighted
\nregression. With this result, we are able to derive a new, computationally
\nefficient FQI algorithm which can even deal with high dimensional action spaces.

关键词

Computer scienceAction (physics)Process (computing)Quality (philosophy)Artificial intelligenceSelection (genetic algorithm)Action selectionGreedy algorithmMachine learningMathematical optimization

相关论文

查看 OTHER 分类全部论文