Jan Leike
Papers
1
Total Citations
508
H-Index
1
About
Jan Leike is a leading researcher in artificial intelligence safety and alignment, whose work has fundamentally shaped how we train advanced AI systems to act in accordance with human values. His most influential contribution, the 2017 paper *Deep Reinforcement Learning from Human Preferences* (508 citations), pioneered a practical method for teaching complex goals to reinforcement learning agents using simple human feedback on trajectory segments—a breakthrough that bypassed the need for explicit reward engineering. This approach has become a cornerstone of modern alignment research, directly influencing techniques used in systems like ChatGPT. Leike’s broader research spans reward modeling, scalable oversight, and the theoretical foundations of AI safety. As a key figure at DeepMind and later OpenAI, he has been instrumental in developing frameworks for aligning superhuman AI, including work on iterative amplification and debate. His impact is measured not only in citations but in the real-world adoption of his methods, making him a pivotal voice in ensuring that increasingly capable AI systems remain beneficial and controllable.
Research Focus
Key Achievements
Top Papers
- 1Deep reinforcement learning from human preferences508 citations · 2017