Yuping Luo
Papers
2
Total Citations
8
H-Index
2
About
Yuping Luo is a researcher specializing in imitation learning and reinforcement learning, with a focus on addressing the covariate shift problem that plagues policy learning from demonstrations. In their most-cited work, "Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling" (2019/2020, 4 citations), Luo introduces a novel framework that combines imitation learning with reinforcement learning to create self-correctable policies. The key innovation lies in using negative sampling—deliberately exposing the policy to off-demonstration states during training—which enables the learned policy to recover from errors and avoid cascading failures common in standard behavioral cloning. This approach bridges the gap between sample-efficient imitation and robust reinforcement learning, offering a practical solution for complex control tasks. While early in their career, Luo’s work demonstrates a clear contribution to making learned policies more resilient in real-world scenarios. Their research is particularly valuable for students and practitioners seeking to build reliable autonomous systems that can learn from limited expert data while maintaining robustness under distribution shift.
Research Focus
Key Achievements
Top Papers
- 1
- 2