Papers
7
Total Citations
81
H-Index
5
About
Shibiao Xu is a dynamic researcher at the forefront of computer vision, robotics, and multimodal artificial intelligence. His work spans 3D scene understanding, human motion generation, robotic grasping, and vision-language integration — areas that collectively address some of the most pressing challenges in embodied AI and autonomous systems. Xu's most impactful contribution, MRFTrans (2024), introduces a novel multimodal representation fusion transformer for monocular 3D semantic scene completion, accumulating 24 citations within its first year — a testament to the method's relevance in spatial perception research. His StableMoFusion framework (2024, 19 citations) advances diffusion-based human motion generation by systematically clarifying architectural design choices, offering a robust and efficient benchmark for the field. Alongside a widely read survey on multimodal fusion and vision-language models for robot vision (2025, 26 combined citations), these works signal Xu's growing influence in bridging perception and language for intelligent systems. His earlier contributions, including 6-DoF robotic grasping in unstructured environments and NeRF-based 6D pose estimation, demonstrate a consistent focus on making robots more capable and perceptually aware in real-world conditions. Xu represents an emerging voice shaping the next generation of intelligent, multimodal robotic systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3Multimodal fusion and vision–language models: A survey for robot vision19 citations · 2025
- 4Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision7 citations · 2025
- 5
- 6VANE-IN: Velocity Auto-Encoder for Inertial Navigation3 citations · 2024
- 7C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation3 citations · 2024