Jintong Li
Papers
1
Total Citations
4
H-Index
1
About
Jintong Li is a rising star in embodied AI and computer vision, whose work tackles the grand challenge of enabling autonomous agents to navigate complex, real-world urban environments. His landmark paper, "CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos" (2025), introduces a paradigm-shifting approach that leverages massive, unlabeled web video data to teach agents how to move through dynamic cityscapes without relying on pre-built maps. This work directly addresses critical limitations in existing navigation systems—such as their failure in off-street or map-free settings—by imbuing agents with common-sense spatial reasoning and the ability to adapt to unpredictable pedestrian flows. Though early in its trajectory, CityWalker has already garnered 4 citations, signaling its potential to reshape autonomous navigation research. Li’s contributions are particularly notable for bridging the gap between large-scale video understanding and embodied decision-making, a frontier that promises to unlock safer, more capable robots and autonomous vehicles. His research sits at the intersection of reinforcement learning, video understanding, and robotics, pushing the boundaries of how machines perceive and act within the messy, rule-bound world of human spaces.
Research Focus
Key Achievements
Top Papers
- 1CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos4 citations · 2025