Home /Research /Primal-Dual Multi-Agent Trust Region Policy Optimization for Safe Multi-Agent Reinforcement Learning
SWARM

Primal-Dual Multi-Agent Trust Region Policy Optimization for Safe Multi-Agent Reinforcement Learning

Jie Li, Junjie Fu

Year
2024
Citations
2

Abstract

In the field of multi-agent reinforcement learning (MARL), achieving high performance is crucial for the success of multi-robot systems. Meanwhile, avoiding unsafe behaviors is becoming a practical problem that needs to be addressed. However, ensuring safety in MARL remains challenging due to the necessity for each agent to not only ensure its own safety but also consider the safety of other agents to ensure overall safe team behavior. In this study, we propose a novel safe multi-agent reinforcement learning algorithm called MATRPO-Lagrangian, which addresses the multi-agent constrained policy optimization problem using a combination of the primal-dual method and trust region policy optimization. Experimental results on the safe MARL benchmark Safe Multi-Agent MuJoCo show that our method achieves competitive performance and significant safe constraint satisfaction ability compared to existing methods.

Keywords

Reinforcement learningDual (grammatical number)Computer scienceMulti-agent systemTrust regionReinforcementArtificial intelligenceComputer securityMaterials scienceComposite material

Related papers

Browse all SWARM papers