Jordi Orbay

Papers

1

Total Citations

3

H-Index

1

About

Jordi Orbay is a researcher whose work sits at the intersection of deep reinforcement learning and scalable AI systems, with a particular focus on how value functions are trained in modern RL architectures. His most notable contribution, the 2024 paper "Stop Regressing: Training Value Functions via Classification for Scalable Deep RL," challenges a foundational assumption in the field. Orbay demonstrates that the traditional mean squared error regression objective used to train value functions can be replaced with a classification-based approach, leading to more stable and scalable learning in large networks. This insight has already garnered early citations and signals a potential paradigm shift in how value-based RL methods are implemented. By rethinking the training objective, Orbay addresses a critical bottleneck in scaling RL to complex, real-world problems. His work is particularly relevant for researchers tackling high-dimensional state spaces and sparse reward environments. With a keen eye for foundational improvements, Orbay is helping to build the next generation of robust, scalable reinforcement learning algorithms that can handle the demands of large-scale deployment.

Research Focus

Key Achievements

1
H-Index
1
Papers
3
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
3 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 11

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago