首页 /研究 /Locally weighted least squares policy iteration for model-free learning in uncertain environments
OTHER

Locally weighted least squares policy iteration for model-free learning in uncertain environments

Matthew Howard, Yoshihiko Nakamura

发表年份
2013
引用次数
4

摘要

This paper introduces Locally Weighted Least Squares Policy Iteration for learning approximate optimal control in settings where models of the dynamics and cost function are either unavailable or hard to obtain. Building on recent advances in Least Squares Temporal Difference Learning, the proposed approach is able to learn from data collected from interactions with a system, in order to build a global control policy based on localised models of the state-action value function. Evaluations are reported characterising learning performance for non-linear control problems including an under-powered pendulum swing-up task, and a robotic door-opening problem under different dynamical conditions.

关键词

Computer scienceMathematical optimizationBellman equationTask (project management)Least-squares function approximationTemporal difference learningSwingFunction (biology)Reinforcement learningState (computer science)

相关论文

查看 OTHER 分类全部论文