Xiaoteng Ma

Tsinghua University

Papers

2

Total Citations

6

H-Index

2

About

Xiaoteng Ma is a leading researcher in reinforcement learning (RL), with a focus on risk-aware decision-making and cross-domain policy generalization. His work addresses critical challenges in deploying RL in high-stakes, real-world environments—from finance and robotics to autonomous driving. Ma’s major contribution includes the development of **Mean-Semivariance Policy Optimization**, a risk-averse RL framework that penalizes only downside volatility, offering a more natural and practical alternative to traditional variance-based risk measures. This work has garnered attention for its ability to balance reward maximization with safety-critical control. He has also pioneered **Value-Guided Data Filtering**, a method for cross-domain policy adaptation that enables agents trained in simulation (source domain) to generalize effectively to real-world environments (target domain) despite dynamics mismatches. With each of his most-cited papers accumulating 3 citations, Ma’s research is recognized for its theoretical rigor and practical relevance. His contributions are shaping the next generation of robust, transferable, and risk-sensitive RL systems, making him a notable figure in the field’s push toward safer and more adaptable AI.

Research Focus

Key Achievements

2
H-Index
2
Papers
6
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning
3 citations · 2022
📈 Most Prolific Year: 2022 (1 Papers)
🤝 Key Collaborators: 10
🏛 Institutions: Tsinghua University

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago