Timo Kaufmann
Papers
2
Total Citations
35
H-Index
1
About
Timo Kaufmann is a leading researcher at the intersection of reinforcement learning and human–AI alignment, with a primary focus on **reinforcement learning from human feedback (RLHF)**. His work tackles one of the field’s most critical challenges: how to train agents to perform complex behaviors without requiring engineers to handcraft reward functions. In his highly cited 2023 survey, *“A Survey of Reinforcement Learning from Human Feedback,”* Kaufmann provides a comprehensive synthesis of RLHF and its predecessor, preference-based reinforcement learning (PbRL), establishing a foundational reference that has already garnered **34 citations** and is widely used by both newcomers and experts. Building on this, his 2025 paper *“DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human Feedback”* introduces a novel framework for intelligently selecting which queries to present to human trainers—addressing a key bottleneck in RLHF scalability. By focusing on diversity, uncertainty, and on-policy sampling, Kaufmann’s DUO method promises to make human feedback more efficient and robust. His contributions are shaping how autonomous systems learn from human guidance, with direct implications for safer, more aligned AI.
Research Focus
Key Achievements
Top Papers
- 1A Survey of Reinforcement Learning from Human Feedback34 citations · 2023
- 2