About

Stefan Elfwing is a leading researcher in reinforcement learning (RL) and evolutionary robotics, whose work bridges the gap between biological decision-making and artificial intelligence. His most significant contributions center on developing safe, modular, and biologically inspired RL architectures. Elfwing pioneered the MaxPain algorithm, which parallelizes reward and punishment prediction—a departure from traditional RL that treats punishments as negative rewards—enabling safer autonomous navigation and control in robots. His work on modular deep RL, which decomposes complex tasks into parallel sub-goals, has been highly influential, with his 2020 paper on the topic garnering 60 citations. Elfwing has also advanced embodied evolution, demonstrating how robots can autonomously evolve survival behaviors without human intervention, and developed free-energy based RL methods for high-dimensional state spaces. His research consistently draws inspiration from neuroscience, particularly the separate neural systems for reward and punishment observed in animals. With over 260 total citations across his top ten papers, Elfwing’s work has shaped modern approaches to safe, efficient, and biologically plausible RL systems, making him a key figure in the field.

Research Focus

Key Achievements

10
H-Index
13
Papers
273
Total Citations
21
Avg Citations/Paper
🏆 Most Cited Paper
Modular deep reinforcement learning from reward and punishment for robot navigation
60 citations · 2020
📈 Most Prolific Year: 2009 (2 Papers)
🤝 Key Collaborators: 7
🏛 Institutions: Contextual Change (United States), Advanced Telecommunications Research Institute International, Okinawa Institute of Science and Technology Graduate University, KTH Royal Institute of Technology, Sandia National Laboratories

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago