首页 /研究 /Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
LEARNING

Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

Johannes Ackermann, Volker Gabler, Takayuki Osa, Masashi Sugiyama

发表年份
2019
引用次数
70
访问权限
开放获取

摘要

Many real world tasks require multiple agents to work together. Multi-agent reinforcement learning (RL) methods have been proposed in recent years to solve these tasks, but current methods often fail to efficiently learn policies. We thus investigate the presence of a common weakness in single-agent RL, namely value function overestimation bias, in the multi-agent setting. Based on our findings, we propose an approach that reduces this bias by using double centralized critics. We evaluate it on six mixed cooperative-competitive tasks, showing a significant advantage over current methods. Finally, we investigate the application of multi-agent methods to high-dimensional robotic tasks and show that our approach can be used to learn decentralized policies in this domain.

关键词

Reinforcement learningComputer scienceDomain (mathematical analysis)Function (biology)Artificial intelligenceMulti-agent systemValue (mathematics)Machine learningMathematics

相关论文

查看 LEARNING 分类全部论文