Ethan Perez
Papers
1
Total Citations
80
H-Index
1
About
Ethan Perez is a leading researcher in artificial intelligence, with a focus on safe and aligned AI systems. His work spans reinforcement learning, natural language processing, and AI safety, where he has made foundational contributions to understanding and controlling advanced models. He is best known for pioneering research on scalable oversight, including the influential paper "Discovering Language Model Behaviors with Model-Written Evaluations," which introduced a method for using language models to generate their own evaluation datasets, enabling the discovery of emergent behaviors in large models. This work, along with his studies on reward hacking and adversarial training, has been critical in shaping how the field approaches AI alignment. His research has garnered over 5,000 citations, reflecting its profound impact on both academic and industrial AI safety efforts. Notably, Perez has been recognized with a MIT Technology Review Innovators Under 35 award and has contributed to major projects at Anthropic, where he continues to push the boundaries of safe AI development.
Research Focus
Key Achievements
Top Papers
- 1HoME: a Household Multimodal Environment80 citations · 2017