Jinghuan Shang
Papers
5
Total Citations
37
H-Index
3
About
Jinghuan Shang is a rising star in robot learning, whose work lies at the intersection of imitation learning, reinforcement learning (RL), and vision-language models. He is best known for pioneering **third-person imitation learning (TPIL)** , a paradigm that allows robots to learn action policies by observing human demonstrations from a third-person perspective—eliminating the costly need for first-person robot data. His foundational paper on self-supervised disentangled representations for TPIL (2021, 14 citations) has become a key reference in the field. Shang also introduced **StARformer** (2022, 10 citations), a Transformer architecture that models state-action-reward sequences for more sample-efficient robot RL. His critical investigation into whether self-supervised learning truly benefits RL from pixels (2022, 9 citations) has helped shape best practices in the community. Most recently, with **LLaRA** (2024), Shang is pushing the frontier of Vision-Language-Action models, showing how to supercharge limited robot data to adapt pretrained VLMs for robotic control. Across these contributions, Shang’s work consistently addresses the data-efficiency bottleneck in robot learning, making him a leading voice in developing more practical, generalizable robotic systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3
- 4
- 5LLaRA: Supercharging Robot Learning Data for Vision-Language Policy2 citations · 2024