Adam Gleave

University of California, Berkeley

Papers

2

Total Citations

107

H-Index

2

About

Adam Gleave is a leading researcher at the intersection of artificial intelligence safety and reinforcement learning (RL), whose work has fundamentally reshaped how we think about robustness in autonomous systems. His most influential contribution is the seminal paper "Adversarial Policies: Attacking Deep Reinforcement Learning" (92 citations), which introduced a critical vulnerability: an RL agent can be defeated by an adversarial policy that manipulates its observations, even without direct access to the victim's inputs. This work exposed a new class of security risks in multi-agent systems, from autonomous driving to game-playing AI. Gleave also advanced the theoretical foundations of reward learning with "Multi-task Maximum Entropy Inverse Reinforcement Learning" (15 citations), developing scalable algorithms to infer multiple reward functions from expert demonstrations—a key challenge for aligning AI with human intent. His research has been instrumental in bridging the gap between adversarial robustness and inverse reinforcement learning, earning him recognition as a rising star in AI safety. By revealing how easily RL agents can be exploited and proposing methods to learn robust reward structures, Gleave’s work provides essential guardrails for deploying AI in high-stakes environments.

Research Focus

Key Achievements

2
H-Index
2
Papers
107
Total Citations
54
Avg Citations/Paper
🏆 Most Cited Paper
Adversarial Policies: Attacking Deep Reinforcement Learning
92 citations · 2019
📈 Most Prolific Year: 2019 (1 Papers)
🤝 Key Collaborators: 6
🏛 Institutions: University of California, Berkeley

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago