Papers

4

Total Citations

192

H-Index

3

About

Zuxuan Wu is a leading researcher at the intersection of computer vision, natural language processing, and embodied AI. His work focuses on enabling intelligent agents to perceive, reason, and act within visual environments by integrating language understanding with decision-making. Wu is best known for pioneering the concept of "The Regretful Agent" (2019, 173 citations), a seminal contribution to Vision and Language Navigation (VLN). This work introduced a heuristic-aided navigation framework that uses progress estimation to allow agents to "regret" past decisions and recover from mistakes, dramatically improving navigation success rates in complex, real-world environments. This innovation set a new standard for embodied AI systems that must follow natural language instructions. Wu has also made significant contributions to large-scale video understanding, co-organizing the LSVC2017 challenge to advance web video classification. Most recently, his work on instruction-guided video prediction (2025) adapts diffusion models to generate future video frames from a single initial frame and a text command, opening new possibilities for content creation and robotics. With over 190 citations, Wu's research continues to shape how machines bridge vision, language, and action.

Research Focus

Key Achievements

3
H-Index
4
Papers
192
Total Citations
48
Avg Citations/Paper
🏆 Most Cited Paper
The Regretful Agent: Heuristic-Aided Navigation Through Progress Estimation
173 citations · 2019
📈 Most Prolific Year: 2019 (2 Papers)
🤝 Key Collaborators: 10
🏛 Institutions: University of Maryland, College Park, Shanghai Key Laboratory of Trustworthy Computing

Top Papers

  1. 1
  2. 2
  3. 3
    LSVC2017
    3 citations · 2017
  4. 4

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago