Xiaoteng Ma
Papers
2
Total Citations
6
H-Index
2
About
Xiaoteng Ma is a leading researcher in reinforcement learning (RL), with a focus on risk-aware decision-making and cross-domain policy generalization. His work addresses critical challenges in deploying RL in high-stakes, real-world environments—from finance and robotics to autonomous driving. Ma’s major contribution includes the development of **Mean-Semivariance Policy Optimization**, a risk-averse RL framework that penalizes only downside volatility, offering a more natural and practical alternative to traditional variance-based risk measures. This work has garnered attention for its ability to balance reward maximization with safety-critical control. He has also pioneered **Value-Guided Data Filtering**, a method for cross-domain policy adaptation that enables agents trained in simulation (source domain) to generalize effectively to real-world environments (target domain) despite dynamics mismatches. With each of his most-cited papers accumulating 3 citations, Ma’s research is recognized for its theoretical rigor and practical relevance. His contributions are shaping the next generation of robust, transferable, and risk-sensitive RL systems, making him a notable figure in the field’s push toward safer and more adaptable AI.
Research Focus
Key Achievements
Top Papers
- 1
- 2Cross-Domain Policy Adaptation via Value-Guided Data Filtering3 citations · 2023