Papers

3

Total Citations

62

H-Index

2

About

Yicong Hong is a leading researcher in embodied AI, specializing in vision-and-language navigation (VLN) and the integration of large vision-language models (VLMs) with robotic reasoning. His work bridges the gap between high-level linguistic instructions and low-level visual perception, enabling agents to navigate unseen environments with unprecedented generalization. Hong’s most influential contribution is **NavGPT-2** (2024, 31 citations), which unlocks navigational reasoning in VLMs by grounding spatial semantics in real-time visual inputs—a breakthrough for sim-to-real transfer. Earlier, his **semantic map supervision** approach (2023, 29 citations) redefined visual representation learning for navigation, teaching agents to encode both object semantics and spatial structure from egocentric video, outperforming traditional classification or self-supervised backbones. Most recently, **NaVid** (2024) pioneers video-based VLM planning for VLN, tackling long-standing challenges in out-of-distribution scene generalization. With over 60 citations across his top papers, Hong’s work is rapidly shaping how embodied agents understand and act in dynamic environments. His research is essential reading for anyone interested in the frontier of language-guided robotics and multimodal reasoning.

Research Focus

Key Achievements

2
H-Index
3
Papers
62
Total Citations
21
Avg Citations/Paper
🏆 Most Cited Paper
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
31 citations · 2024
📈 Most Prolific Year: 2024 (2 Papers)
🤝 Key Collaborators: 17
🏛 Institutions: Adobe Systems (United States), Australian National University

Top Papers

  1. 1
  2. 2
  3. 3

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago