Papers
5
Total Citations
53
H-Index
3
About
Rongtao Xu is an emerging researcher at the forefront of embodied AI, multimodal learning, and 3D scene understanding. His work spans several interconnected domains, including vision-language navigation, multimodal representation fusion, and neural rendering-based pose estimation — areas that collectively push the boundaries of how intelligent agents perceive and interact with complex environments. Among his most notable contributions is MRFTrans, a Multimodal Representation Fusion Transformer designed for monocular 3D semantic scene completion, which has already garnered 24 citations since its 2024 publication, signaling strong community interest. His comprehensive survey on multimodal fusion and vision-language models for robot vision (19 citations) has quickly become a valuable reference for researchers navigating this rapidly evolving landscape. Xu has also tackled the particularly challenging problem of zero-shot Vision-Language Navigation in Continuous Environments, developing constraint-aware methods that eliminate the need for expert demonstrations. Additional work includes NaVid, a video-based approach improving generalization in sim-to-real navigation transfer, and C2Fi-NeRF, a coarse-to-fine neural radiance field method for 6D pose estimation. Across his portfolio, Xu demonstrates a consistent commitment to bridging perception, language grounding, and real-world robotic applicability.
Research Focus
Key Achievements
Top Papers
- 1
- 2Multimodal fusion and vision–language models: A survey for robot vision19 citations · 2025
- 3
- 4C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation3 citations · 2024
- 5