Papers
16
Total Citations
558
H-Index
9
About
Dylan Hadfield-Menell is a pioneering researcher at the intersection of artificial intelligence, robotics, and AI safety, whose work has fundamentally shaped how we think about aligning autonomous systems with human values. He is perhaps best known for introducing Cooperative Inverse Reinforcement Learning (CIRL) in 2016, a landmark framework — now with over 326 citations — that formally defines the value alignment problem by modeling the relationship between humans and AI as a cooperative game. This foundational contribution has become a cornerstone of AI safety research worldwide. His broader research agenda tackles some of the most pressing challenges in human-robot interaction and safe AI design. Through works like "The Off-Switch Game" and "Should Robots be Obedient?", he explores the nuanced dynamics of human oversight, demonstrating why well-intentioned obedience or self-preservation in AI systems can paradoxically undermine human welfare. His investigations into preference learning, including "The Assistive Multi-Armed Bandit," grapple with the reality that humans themselves are imperfect decision-makers. Earlier in his career, Hadfield-Menell contributed meaningfully to robotic manipulation and task planning under uncertainty. Collectively, his work offers a rigorous, mathematically grounded vision for building AI systems that are genuinely beneficial — making him an essential voice in contemporary AI safety discourse.
Research Focus
Key Achievements
Top Papers
- 1Cooperative Inverse Reinforcement Learning326 citations · 2016
- 2Modular task and motion planning in belief space42 citations · 2015
- 3On the Utility of Model Learning in HRI42 citations · 2019
- 4
- 5The Assistive Multi-Armed Bandit31 citations · 2019
- 6Should Robots be Obedient?19 citations · 2017
- 7Pragmatic-Pedagogic Value Alignment16 citations · 2019
- 8The Off-Switch Game12 citations · 2017
- 9On the Utility of Model Learning in HRI11 citations · 2019
- 10Beyond lowest-warping cost action selection in trajectory transfer4 citations · 2015