Papers
3
Total Citations
62
H-Index
2
About
Yicong Hong is a leading researcher in embodied AI, specializing in vision-and-language navigation (VLN) and the integration of large vision-language models (VLMs) with robotic reasoning. His work bridges the gap between high-level linguistic instructions and low-level visual perception, enabling agents to navigate unseen environments with unprecedented generalization. Hong’s most influential contribution is **NavGPT-2** (2024, 31 citations), which unlocks navigational reasoning in VLMs by grounding spatial semantics in real-time visual inputs—a breakthrough for sim-to-real transfer. Earlier, his **semantic map supervision** approach (2023, 29 citations) redefined visual representation learning for navigation, teaching agents to encode both object semantics and spatial structure from egocentric video, outperforming traditional classification or self-supervised backbones. Most recently, **NaVid** (2024) pioneers video-based VLM planning for VLN, tackling long-standing challenges in out-of-distribution scene generalization. With over 60 citations across his top papers, Hong’s work is rapidly shaping how embodied agents understand and act in dynamic environments. His research is essential reading for anyone interested in the frontier of language-guided robotics and multimodal reasoning.
Research Focus
Key Achievements
Top Papers
- 1
- 2Learning Navigational Visual Representations with Semantic Map Supervision29 citations · 2023
- 3