Jordi Orbay
Papers
1
Total Citations
3
H-Index
1
About
Jordi Orbay is a researcher whose work sits at the intersection of deep reinforcement learning and scalable AI systems, with a particular focus on how value functions are trained in modern RL architectures. His most notable contribution, the 2024 paper "Stop Regressing: Training Value Functions via Classification for Scalable Deep RL," challenges a foundational assumption in the field. Orbay demonstrates that the traditional mean squared error regression objective used to train value functions can be replaced with a classification-based approach, leading to more stable and scalable learning in large networks. This insight has already garnered early citations and signals a potential paradigm shift in how value-based RL methods are implemented. By rethinking the training objective, Orbay addresses a critical bottleneck in scaling RL to complex, real-world problems. His work is particularly relevant for researchers tackling high-dimensional state spaces and sparse reward environments. With a keen eye for foundational improvements, Orbay is helping to build the next generation of robust, scalable reinforcement learning algorithms that can handle the demands of large-scale deployment.
Research Focus
Key Achievements
Top Papers
- 1