首页 /研究 /Variable risk control via stochastic optimization
MANIPULATION

Variable risk control via stochastic optimization

Scott Kuindersma, Roderic A. Grupen, Andrew G. Barto

发表年份
2013
引用次数
37

摘要

We present new global and local policy search algorithms suitable for problems with policy-dependent cost variance (or risk), a property present in many robot control tasks. These algorithms exploit new techniques in non-parametric heteroscedastic regression to directly model the policy-dependent distribution of cost. For local search, the learned cost model can be used as a critic for performing risk-sensitive gradient descent. Alternatively, decision-theoretic criteria can be applied to globally select policies to balance exploration and exploitation in a principled way, or to perform greedy minimization with respect to various risk-sensitive criteria. This separation of learning and policy selection permits variable risk control, where risk-sensitivity can be flexibly adjusted and appropriate policies can be selected at runtime without relearning. We describe experiments in dynamic stabilization and manipulation with a mobile manipulator that demonstrate learning of flexible, risk-sensitive policies in very few trials.

关键词

Computer scienceExploitMathematical optimizationControl variableParametric statisticsControl (management)Variance (accounting)Sensitivity (control systems)Machine learningArtificial intelligence

相关论文

查看 MANIPULATION 分类全部论文