Donglai Xiang
Papers
2
Total Citations
26
H-Index
2
About
Donglai Xiang is a rising star in embodied AI and computer vision, whose research bridges the gap between human understanding and robotic action. His primary contributions lie in two intersecting domains: **vision-language-action models (VLAs)** for generalizable robot control, and **egocentric body pose estimation** for human-centric computing. In his seminal 2025 work, "CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models" (23 citations), Xiang pioneered a novel framework that integrates visual chain-of-thought reasoning into VLAs, enabling robots to leverage both robotic and non-robotic data for more robust sensorimotor control—a critical step toward general-purpose robotics. Earlier, his 2020 paper "You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions" (3 citations) introduced a learning-based method to estimate the camera wearer's 3D body pose from egocentric video, a challenging problem with applications in augmented reality and healthcare. By modeling the interaction between the first-person (wearer) and second-person (observed) views, Xiang’s work has laid foundational insights for understanding human motion from wearable cameras. His research continues to push boundaries in making AI systems more perceptive and actionable in real-world environments.
Research Focus
Key Achievements
Top Papers
- 1CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models23 citations · 2025
- 2