Andrew Critch
Papers
2
Total Citations
11
H-Index
2
About
Andrew Critch is a leading researcher at the intersection of artificial intelligence safety, game theory, and human-AI alignment. His work critically examines how AI systems can infer and act upon human preferences, especially when those preferences are irrational or conflicting. In his highly cited 2021 paper, *"Human irrationality: both bad and good for reward inference,"* Critch explores how deviations from rational behavior—often dismissed as noise—can actually provide valuable signals for reward learning. This work has become foundational for researchers designing AI that robustly interprets human goals. Critch also made a landmark contribution to multi-agent alignment with his 2020 paper, *"Multi-Principal Assistance Games: Definition and Collegial Mechanisms."* Here, he introduced a novel framework for a single AI to assist multiple human principals with divergent values, cleverly circumventing classic impossibility results like Gibbard's theorem. This work has garnered attention for its practical approach to democratic AI governance. Beyond his publications, Critch is a co-founder of the AI safety organization *Encultured AI* and a former research scientist at DeepMind, where he helped shape the field's understanding of scalable oversight and value alignment.
Research Focus
Key Achievements
Top Papers
- 1Human irrationality: both bad and good for reward inference7 citations · 2021
- 2Multi-Principal Assistance Games: Definition and Collegial Mechanisms4 citations · 2020