首页 /研究 /Primal-Dual Multi-Agent Trust Region Policy Optimization for Safe Multi-Agent Reinforcement Learning
SWARM

Primal-Dual Multi-Agent Trust Region Policy Optimization for Safe Multi-Agent Reinforcement Learning

Jie Li, Junjie Fu

发表年份
2024
引用次数
2

摘要

In the field of multi-agent reinforcement learning (MARL), achieving high performance is crucial for the success of multi-robot systems. Meanwhile, avoiding unsafe behaviors is becoming a practical problem that needs to be addressed. However, ensuring safety in MARL remains challenging due to the necessity for each agent to not only ensure its own safety but also consider the safety of other agents to ensure overall safe team behavior. In this study, we propose a novel safe multi-agent reinforcement learning algorithm called MATRPO-Lagrangian, which addresses the multi-agent constrained policy optimization problem using a combination of the primal-dual method and trust region policy optimization. Experimental results on the safe MARL benchmark Safe Multi-Agent MuJoCo show that our method achieves competitive performance and significant safe constraint satisfaction ability compared to existing methods.

关键词

Reinforcement learningDual (grammatical number)Computer scienceMulti-agent systemTrust regionReinforcementArtificial intelligenceComputer securityMaterials scienceComposite material

相关论文

查看 SWARM 分类全部论文