Hongtao Wu
Papers
4
Total Citations
35
H-Index
4
About
Hongtao Wu is a researcher at the forefront of embodied AI, specializing in vision-language models, generative pre-training, and robot manipulation. His work addresses one of robotics' most pressing challenges: enabling robots to understand and interact with complex environments through multimodal learning. Wu's most influential contribution, "Vision-Language Foundation Models as Effective Robot Imitators" (19 citations), demonstrated how existing vision-language models can be efficiently fine-tuned to serve as capable robot imitators, lowering the barrier to deploying powerful AI in physical systems. Building on this foundation, his GR-2 framework pushed the boundaries further by pre-training a generalist robot agent on an unprecedented 38 million video clips and over 50 billion tokens sourced from the internet, achieving state-of-the-art versatility in robot manipulation tasks. His earlier work on large-scale video generative pre-training similarly established that visual representations learned from video data transfer effectively to robotic control. More recently, Wu has extended his expertise to legged locomotion, applying world model-based perception to help robots navigate challenging terrains. Collectively, his research demonstrates a consistent and impactful vision: harnessing internet-scale data and generative modeling to build more capable, generalizable robotic systems.
Research Focus
Key Achievements
Top Papers
- 1Vision-Language Foundation Models as Effective Robot Imitators19 citations · 2023
- 2
- 3World Model-Based Perception for Visual Legged Locomotion5 citations · 2025
- 4