Adam Gleave
Papers
2
Total Citations
107
H-Index
2
About
Adam Gleave is a leading researcher at the intersection of artificial intelligence safety and reinforcement learning (RL), whose work has fundamentally reshaped how we think about robustness in autonomous systems. His most influential contribution is the seminal paper "Adversarial Policies: Attacking Deep Reinforcement Learning" (92 citations), which introduced a critical vulnerability: an RL agent can be defeated by an adversarial policy that manipulates its observations, even without direct access to the victim's inputs. This work exposed a new class of security risks in multi-agent systems, from autonomous driving to game-playing AI. Gleave also advanced the theoretical foundations of reward learning with "Multi-task Maximum Entropy Inverse Reinforcement Learning" (15 citations), developing scalable algorithms to infer multiple reward functions from expert demonstrations—a key challenge for aligning AI with human intent. His research has been instrumental in bridging the gap between adversarial robustness and inverse reinforcement learning, earning him recognition as a rising star in AI safety. By revealing how easily RL agents can be exploited and proposing methods to learn robust reward structures, Gleave’s work provides essential guardrails for deploying AI in high-stakes environments.
Research Focus
Key Achievements
Top Papers
- 1Adversarial Policies: Attacking Deep Reinforcement Learning92 citations · 2019
- 2Multi-task Maximum Entropy Inverse Reinforcement Learning15 citations · 2018