Mohammad Ghavamzadeh
University of Alberta, Adobe Systems (United States), Meta (Israel)
Papers
5
Total Citations
428
H-Index
5
About
Mohammad Ghavamzadeh is a prominent researcher whose work spans safe reinforcement learning, hierarchical multi-agent systems, and constrained decision-making. He has made foundational contributions to the theory and practice of deploying intelligent agents in real-world settings where both performance and safety are paramount. Ghavamzadeh's most influential contributions center on Lyapunov-based approaches to safe reinforcement learning, developing rigorous mathematical frameworks that guarantee agents avoid dangerous states during both training and deployment. His 2019 paper on safe policy optimization for continuous control, with 153 citations, introduced constrained Markov decision process formulations that have become a cornerstone reference in the safe RL community. An earlier 2018 companion work further established Lyapunov stability theory as a practical tool for enforcing safety constraints in RL. Beyond safety, Ghavamzadeh has significantly advanced multi-agent reinforcement learning through hierarchical frameworks. His 2004 and 2006 papers introduced cooperative hierarchical multi-agent RL algorithms that accelerate learning in complex team tasks, collectively accumulating nearly 200 citations and remaining highly relevant to modern multi-agent research. His more recent work on conservative exploration in bandits reflects a consistent intellectual thread: enabling learning systems to improve over existing policies while maintaining reliable, risk-aware behavior — a challenge of growing importance in healthcare, robotics, and digital applications.
Research Focus
Key Achievements
Top Papers
- 1Lyapunov-based Safe Policy Optimization for Continuous Control153 citations · 2019
- 2Hierarchical multi-agent reinforcement learning129 citations · 2006
- 3A Lyapunov-based Approach to Safe Reinforcement Learning78 citations · 2018
- 4Hierarchical Multiagent Reinforcement Learning56 citations · 2004
- 5Improved Algorithms for Conservative Exploration in Bandits12 citations · 2020