Timo Kaufmann

Ludwig-Maximilians-Universität München

Papers

2

Total Citations

35

H-Index

1

About

Timo Kaufmann is a leading researcher at the intersection of reinforcement learning and human–AI alignment, with a primary focus on **reinforcement learning from human feedback (RLHF)**. His work tackles one of the field’s most critical challenges: how to train agents to perform complex behaviors without requiring engineers to handcraft reward functions. In his highly cited 2023 survey, *“A Survey of Reinforcement Learning from Human Feedback,”* Kaufmann provides a comprehensive synthesis of RLHF and its predecessor, preference-based reinforcement learning (PbRL), establishing a foundational reference that has already garnered **34 citations** and is widely used by both newcomers and experts. Building on this, his 2025 paper *“DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human Feedback”* introduces a novel framework for intelligently selecting which queries to present to human trainers—addressing a key bottleneck in RLHF scalability. By focusing on diversity, uncertainty, and on-policy sampling, Kaufmann’s DUO method promises to make human feedback more efficient and robust. His contributions are shaping how autonomous systems learn from human guidance, with direct implications for safer, more aligned AI.

Research Focus

Key Achievements

1
H-Index
2
Papers
35
Total Citations
18
Avg Citations/Paper
🏆 Most Cited Paper
A Survey of Reinforcement Learning from Human Feedback
34 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 7
🏛 Institutions: Ludwig-Maximilians-Universität München

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago